[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-parakeet-redux-ternary-cpu-asr":3,"topics-all":38,"news-related-f3b52d77-bfad-4289-adab-777d19a79bcb":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"f3b52d77-bfad-4289-adab-777d19a79bcb","Parakeet Redux把ASR压到178MB:纯CPU跑出113倍实时","Moondream 发布 parakeet-redux:NVIDIA Parakeet 0.6B 编码器权重三元化(-1\u002F0\u002F+1),178MB 即可部署。Photon 在 8 核 x86 CPU 跑出 113 倍实时,为最快 CPU 运行时 2.5 倍;25 语种 FLEURS 与长音频反超原版,强噪声下退化。","语音转写的默认假设一直是要 GPU:模型动辄上 GB,推理吞吐靠显卡堆。Moondream 最近发布了 [parakeet-redux](https:\u002F\u002Fhuggingface.co\u002Fmoondream\u002Fparakeet-redux),把这个假设拆了——NVIDIA 开源 ASR 模型 parakeet-tdt-0.6b-v3 被压到 178MB,在一台普通 x86 CPU 上跑出 113 倍实时的转写速度,全程不碰显卡。\n\n## 三元化:最激进的量化\n\nparakeet-redux 的做法是对 parakeet-tdt-0.6b-v3 做 1.58-bit 三元量化:架构不动、tokenizer 不动,但编码器的每个权重只剩 -1、0、+1 三个取值。原始权重 1.2GB,压完后 178MB,不到原来的 15%,小到可以塞进任何边缘设备。配套的 Photon 运行时直接读取打包后的权重:x86 上走 AVX-512 VNNI 指令,ARM 上走 NEON,Apple GPU 上走 Metal。\n\n## 速度:同硬件上最快的 Parakeet CPU 方案\n\nMoondream 在 AMD EPYC 9575F(Zen 5)的 8 个物理核上做了同机对比:parakeet-redux 配 Photon 跑到 113 倍实时,是实测最快的其他 Parakeet CPU 运行时的 2.5 倍——parakeet.cpp 的 q8_0 版只有 45 倍,sherpa-onnx 42 倍,onnx-asr 28 倍。Apple M2 上同样领先:CPU 38 倍、GPU 43 倍,而 parakeet.cpp 的 Metal 路径 38 倍、跑 fp32 的 parakeet-mlx 只有 37 倍。一个三元权重模型,在 MacBook Air 上反超了专门适配过的原生运行时。\n\n## 精度:多语种与长音频反超,噪声是短板\n\n英文 Open ASR Leaderboard 七个测试集平均 WER,Redux 6.55 对原版 6.26,差距在 0.3 以内;AMI 会议、Earnings-22 财报电话两集还略有胜出。真正反直觉的是 25 语种 FLEURS:平均 10.56,好于原版的 11.62——爱沙尼亚语 9.15 对 13.23,拉脱维亚语 12.80 对 21.38,斯洛文尼亚语 16.21 对 21.76。10-20 分钟的 TED-LIUM 长音频同样反超,2.51 对 2.71。短板在噪声:MUSAN 九个强噪条件平均 9.04,明显落后原版的 6.72。模型卡自己的解释是三元编码器的声学余量更薄,低信噪比下更容易把词替换成近音词。\n\n## 所以呢\n\n对做本地语音接口、隐私敏感转写的团队,这是一条明确的成本路线:178MB 的模型加一个运行时,笔记本和边缘盒子上就能跑实时转写,不必把音频送上云。两点提醒:其一,Moondream 采用自定义许可证,商用前要读条款;其二,强噪声场景的退化说明三元化不是免费的,选型时先拿自己的真实音频压测。ASR 生态的去 GPU 化正在成为一条独立赛道——量化到三元、重写运行时、砍掉 Python 依赖,路径各异,方向一致:把推理成本从云端搬到终端。","https:\u002F\u002Fhuggingface.co\u002Fmoondream\u002Fparakeet-redux","62158e77-3507-4013-9d85-9ebde1444211",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c4455801-0273-477d-9e92-00b4addfa439","asr",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1a445691-b232-4827-9d76-8673770762bc","en","Parakeet Redux: ternary ASR at 178MB runs 113x realtime on CPU","Moondream parakeet-redux quantizes NVIDIA 0.6B ASR to ternary: 178MB, 113x realtime on 8 CPU cores, beats the original on FLEURS, trails in noise.","Speech-to-text has always assumed a GPU: models run over a gigabyte, and throughput is bought with graphics cards. Moondream's [parakeet-redux](https:\u002F\u002Fhuggingface.co\u002Fmoondream\u002Fparakeet-redux) breaks that assumption — NVIDIA's open ASR model parakeet-tdt-0.6b-v3 compressed to 178MB, transcribing at 113x realtime on an ordinary x86 CPU without touching a GPU.\n\n## Ternary weights: the most aggressive quantization\n\nparakeet-redux applies 1.58-bit ternary quantization to parakeet-tdt-0.6b-v3: same architecture, same tokenizer, but every encoder weight is now -1, 0, or +1. Weights drop from 1.2GB to 178MB — under 15% of the original, small enough for any edge device. The companion Photon runtime reads the packed weights directly: AVX-512 VNNI on x86, NEON on ARM, Metal on Apple GPUs.\n\n## Speed: the fastest Parakeet CPU setup on the same hardware\n\nOn 8 physical cores of an AMD EPYC 9575F (Zen 5), parakeet-redux with Photon reaches 113x realtime — 2.5x the fastest other Parakeet CPU runtime measured. parakeet.cpp at q8_0 manages 45x, sherpa-onnx 42x, onnx-asr 28x. On an Apple M2 it also leads: 38x on CPU and 43x on GPU, versus 38x for parakeet.cpp's Metal path and 37x for fp32 parakeet-mlx. A ternary-weight model outpaces natively-tuned runtimes on a MacBook Air.\n\n## Accuracy: multilingual and long-form wins, noise is the weak flank\n\nOn the seven English test sets of the Open ASR Leaderboard, Redux averages 6.55 WER versus 6.26 for the original — within 0.3, with AMI and Earnings-22 slightly ahead. The counterintuitive result is 25-language FLEURS: 10.56 average, beating the original's 11.62 — Estonian 9.15 vs 13.23, Latvian 12.80 vs 21.38, Slovene 16.21 vs 21.76. TED-LIUM long-form (10-20 minute talks) also flips: 2.51 vs 2.71. The weak flank is noise: 9.04 average across nine MUSAN conditions, well behind the original's 6.72. The model card's own explanation: the ternary encoder has a thinner acoustic margin, substituting similar-sounding words more often at low SNR.\n\n## So what\n\nFor teams building local voice interfaces or privacy-sensitive transcription, this is a clear cost path: a 178MB model plus one runtime gives realtime transcription on laptops and edge boxes, with no cloud round-trip. Two caveats: Moondream ships it under a custom license — read the terms before commercial use — and the noise regression means ternary quantization is not free; stress-test with your own real audio. The de-GPU-ification of ASR is becoming its own lane: ternary quantization, rewritten runtimes, stripped Python dependencies — different methods, one direction: moving inference cost from the cloud to the endpoint.","parakeet-redux-ternary-cpu-asr","2026-10-01T21:08:38Z","2026-10-01T21:08:41.528288Z","2026-10-01T21:08:41.528298Z",true,"agent",65,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"3e28cfaf-74aa-4521-8eba-37332fe93901","7块ESP32跑1.58-bit Qwen:功耗1.53瓦","esp32s3-bitnet-llm-cluster","2026-09-30T13:11:01+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"6d60f9ba-4866-4520-9f12-d955e37f8472","Gated DeltaNet 全压 4-bit 没掉点:一篇论文拆掉混合 LLM 的量化禁忌","gated-deltanet-nvfp4-full-4bit","2026-09-04T15:08:02+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"b2d34fba-93ee-469b-9f64-6a7568f89745","Liquid AI 用量化感知蒸馏,把 LFM2.5 4-bit 精度拉回 97%","lfm25-qad-quantization-aware-distillation-edge","2026-08-20T11:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"c197245e-0a8a-4028-9ad2-6547bdc01be5","BitNet 团队把 1.58-bit 量化推进到 embedding:检索向量也可以\"训练时就压\"","bitnet-1-58bit-embedding","2026-07-18T03:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"6452bb36-79c8-471b-aa4e-fad99bca9b04","SharQ 用「稀疏-稠密双轨」把 FP4 推理提速 2.4 倍:训练免费还跨平台","sharq-sparse-dense-fp4","2026-07-01T00:00:00+00:00"]