[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-phonon-2-2-bit-parakeet-asr":3,"topics-all":38,"news-related-32aa44d5-81f1-4717-8d17-ff713101b725":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"32aa44d5-81f1-4717-8d17-ff713101b725","Phonon-2:2.1比特量化ASR,164MB逼近全精度","Fermion Research 开源 164MB 英文 ASR 模型 Phonon-2：基于 NVIDIA Parakeet TDT 0.6B v3 量化，编码器每权重约 2.1 比特。七套英文基准平均词错率 5.21%，两集反超 2.5GB 全精度教师，体积缩至 1\u002F15。","164MB 下载、平均词错率 5.21%——这是 Fermion Research 新开源的英文语音识别模型 Phonon-2 交出的成绩单。它的底座不是从零训练，而是 NVIDIA 的 Parakeet TDT 0.6B v3，一个 2.5GB 的全精度模型。把 15 倍的体积差压到接近无损，靠的是一套五电平量化方案：编码器每个权重只保留约 2.1 比特的信息。\n\n## 七个基准集，逐项对数字\n\n按 [官方模型卡](https:\u002F\u002Fhuggingface.co\u002FFermionResearch\u002FPhonon-2) 公布的 Open ASR Leaderboard 测试结果，七个英文集合上 Phonon-2 平均词错率 5.21%，全精度教师为 4.96%，平均差距 0.25 个百分点。分集合看互有胜负：会议语料 AMI 上 9.37 对 9.42、议会演讲 VoxPopuli 上 2.46 对 3.19，这两集反超教师；其余五集（LS clean、LS other、Earnings-22、GigaSpeech、SPGISpeech）教师仍然领先。模型卡自述在议会语音上达到教师词准确率的 100.8%，会议集上直接胜出——与逐集数字一致。\n\n同一张表里还有同赛道的 Parakeet Redux（moondream 出品，三元量化，178MB）：平均 5.21% 对 5.69%，Phonon-2 占优；七个集合中除 AMI（9.37 对 9.16）外逐集领先。两条独立的压缩路线在 200MB 以下这个体积档位正面相遇。\n\n## 2.1 比特怎么存\n\n「五电平」指每个权重被映射到五个可学习层级之一，平均约 2.1 比特，编解码实现就放在仓库的 quint5_codec.py 里。模型卡称其为「900MB 以下最准的开源英文语音识别模型，所有分数更好的开源模型至少是它 5.8 倍大」——这是官方自报口径，暂无第三方复现。速度侧的官方数字：一小时音频在 M5 MacBook Air 上约 20 秒转写完（174 倍实时），八个 Zen 5 核心 CPU 上 143 倍，单张 H100 按 128 路批量跑出 6680 倍。\n\n## 不止是权重文件\n\n[GitHub 仓库](https:\u002F\u002Fgithub.com\u002Ffermionresearch\u002Fphonon) 给出了完整引擎矩阵：Apple Silicon 走 MLX，Linux 与 Windows CPU、NVIDIA CUDA 各有官方 Docker 镜像；命令行支持文件转写和麦克风实时转写，还能起一个 OpenAI 兼容的推理端点。家族产品线上，Phonon-1（415MB）基于 Qwen3-ASR-0.6B，以 Apache-2.0 发布；Phonon-2 权重随上游 Parakeet 采用 CC-BY-4.0，改动清单写在 NOTICE 文件里。它同时是 Mac 听写应用 Detta 的内置引擎——商业闭环和开源权重并不冲突。\n\n## 端侧 ASR 的军备竞赛\n\n一周之内，Parakeet Redux 和 Phonon-2 先后把 Parakeet 系底座压进 200MB 以内：前者赌三元量化的通用性，后者赌五电平量化的精度上限。Phonon-2 真正的看点不在某个单项分数，而在于 2.1 比特的极限压缩做到了单集反超全精度教师——当量化不再自动等于精度天花板，「音频不出本机」就从隐私妥协变成了工程默认选项。\n\n所以呢：下次给 Mac 应用挑语音转写引擎，先看看 164MB 能买到什么，再决定要不要为云 API 付那笔账单。","https:\u002F\u002Fhuggingface.co\u002FFermionResearch\u002FPhonon-2","c0a08179-3527-4f82-a2ff-0d45ad44a8c8",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"c4455801-0273-477d-9e92-00b4addfa439","asr",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b49648f9-963e-4082-8684-3d085b7358fe","quantization",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"a832e8d7-2874-4534-a635-8e122c7bb4c9","en","Phonon-2: 2.1-bit ASR in 164MB, near FP accuracy","Fermion Research's Phonon-2: English ASR in 164MB at ~2.1 bits\u002Fweight, 5.21% WER over seven Open ASR sets, beating its 2.5GB Parakeet teacher on two.","A 164MB download averaging 5.21% word error rate — that is the scorecard Fermion Research published for Phonon-2, its newly open-sourced English speech recognition model. The base is not trained from scratch: it is NVIDIA's Parakeet TDT 0.6B v3, a 2.5GB full-precision model. Closing a 15x size gap to near-lossless comes down to one design choice: a five-level quantization scheme that keeps each encoder weight at roughly 2.1 bits of information.\n\n## Seven benchmarks, digit by digit\n\nPer the [official model card](https:\u002F\u002Fhuggingface.co\u002FFermionResearch\u002FPhonon-2), across the Open ASR Leaderboard's seven English sets Phonon-2 averages 5.21% WER against the full-precision teacher's 4.96% — an average gap of 0.25 points. Set by set, wins go both ways: on the AMI meeting corpus it scores 9.37 versus 9.42, and on parliamentary speech (VoxPopuli) 2.46 versus 3.19, both beating the teacher; on the other five sets (LS clean, LS other, Earnings-22, GigaSpeech, SPGISpeech) the teacher still leads. The card states the model reaches 100.8% of the teacher's word accuracy on parliamentary speech and wins outright on meetings — consistent with the per-set numbers.\n\nThe same table includes a direct rival built on the same foundation: Parakeet Redux from moondream, a ternary-quantized 178MB model. Phonon-2 leads on average, 5.21% versus 5.69%, and set by set everywhere except AMI (9.37 versus 9.16). Two independent compression routes now meet head-on in the sub-200MB weight class.\n\n## How you store 2.1 bits\n\n\"Five levels\" means every weight maps to one of five learned levels, about 2.1 bits on average; the codec implementation sits in the repo's quint5_codec.py. The model card calls it the most accurate open English speech recognition model under 900MB, with every open model scoring better being at least 5.8 times its size — an official, self-reported claim with no third-party replication yet. On speed, the published numbers: an hour of audio transcribed in about 20 seconds on an M5 MacBook Air (174x realtime), 143x on eight Zen 5 CPU cores, and 6,680x on a single H100 at batch 128.\n\n## More than a weights file\n\nThe [GitHub repository](https:\u002F\u002Fgithub.com\u002Ffermionresearch\u002Fphonon) ships a complete engine matrix: MLX on Apple Silicon, official Docker images for Linux and Windows CPUs plus NVIDIA CUDA; the command line handles file transcription and live microphone capture, and can serve an OpenAI-compatible endpoint. On the family tree, Phonon-1 (415MB) is built on Qwen3-ASR-0.6B and released under Apache-2.0, while Phonon-2's weights follow the upstream Parakeet license, CC-BY-4.0, with changes listed in a NOTICE file. Phonon-2 is also the engine inside Detta, a Mac dictation app — commercial closure and open weights are not in conflict.\n\n## The on-device ASR arms race\n\nWithin a single week, Parakeet Redux and Phonon-2 both squeezed a Parakeet foundation under 200MB: one bets on the generality of ternary quantization, the other on the accuracy ceiling of five-level quantization. The real story of Phonon-2 is not any single benchmark number — it is that 2.1-bit extreme quantization managed to beat a full-precision teacher on individual sets. Once quantization no longer automatically means an accuracy ceiling, keeping audio on the machine stops being a privacy compromise and becomes the engineering default.\n\nSo: next time you pick a transcription engine for a Mac app, check what 164MB buys you before paying that cloud API bill.","phonon-2-2-bit-parakeet-asr","2026-10-03T13:10:31Z","2026-10-03T13:10:46.799365Z","2026-10-03T13:10:46.799378Z",true,"agent",255,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"4ccc491b-dbee-4c84-beb6-7cf519f76320","LoRA 基座换 GGUF:40G 显存训 125B","lora-over-gguf-low-vram-training","2026-10-10T21:08:22+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"f3b52d77-bfad-4289-adab-777d19a79bcb","Parakeet Redux把ASR压到178MB:纯CPU跑出113倍实时","parakeet-redux-ternary-cpu-asr","2026-10-01T21:08:38+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"c23c2fca-29cc-40e8-9ff2-aebb32aba767","3.21 比特量化:27B 模型从 54GB 压到 12.3GB","orcasaq2-3-bit-qwen-quantization","2026-09-30T21:20:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"ce40a7fa-acca-4609-a82e-5800f2e1026a","35B 压成 9.96GB 单文件:BTL-4 的 2.30 bit 极限量化和两条反直觉结论","btl-4-compact-9gb-2bit-quantization","2026-08-17T13:20:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"031e715e-c9d6-4855-83da-0515f33f0e3c","POCKET：35B MoE 1-bit 跑进 iPhone，27 tok\u002Fs","pocket-35b-moe-iphone-edge","2026-07-28T04:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"7fd694d4-7eec-4dd4-8fb7-b6e547da98f4","Falcon ASR 开源:1.6B 刷新阿语 WER","falcon-asr-arabic-speech-recognition","2026-10-11T23:08:05+00:00"]