[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-interfaze-diffusion-gemma-asr":3,"news-related-1fe65909-67b4-4019-89fe-76eb4b226c41":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1fe65909-67b4-4019-89fe-76eb4b226c41","Interfaze 把扩散解码塞进 ASR：diffusion-gemma-asr-small 用 42M 适配器撬开六语种开放语音识别","YC 出身的 Interfaze 7 月 2 日开源 `diffusion-gemma-asr-small`,被官方称为首个多语种开源扩散 ASR。可训练参数仅 4200 万,是 26B DiffusionGemma 主干的 0.16%——架构思路\"冻结一切,只训极小适配器\":whisper-small 编码器把 30 秒音频压成 188 个 audio token,scattering 进 DiffusionGemma 的 \u003C|audio|> 槽位,LoRA 让主干注意新模态,扩散解码器在 192 token 画布上做 16 步双向去噪。\n\n工程关键是 CTC 预热:起初直接喂音频给冻结 LLM,loss 卡 8 不动——注意力学会\"忽略噪声\"。修复方案是用 CTC loss 把 audio token 强映射到转写,300 步内 CTC 从 24 掉到 8.6,LibriSpeech test-clean 英文 WER 从 90% 压到 6.6%,反超 Whisfusion(8.3%)和 TransFusion proof-of-concept。\n\n相对自回归 Whisper-small(~3.4%)仍有 3-4 个百分点差距,团队归因数据量而非架构。42M 适配器覆盖英、德、法、西、印地、中六语,FLEURS 英文 WER 15.7%、Mandarin CER 29.6%。扩散 ASR 推理开销由去噪步数决定、与音频长度解耦,8 步 14.9× 实时,适合批量转写。","https:\u002F\u002Fwww.marktechpost.com\u002F2026\u002F07\u002F02\u002Finterfaze-ships-diffusion-gemma-asr-small-an-open-source-diffusion-asr-model-transcribing-six-languages-via-diffusiongemmas-parallel-denoising-decoder\u002F","8382d60c-c2c4-49c5-9638-8518b803f88f",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"600cfbcf-83cc-4c60-9539-7e6314aaf128","en","Interfaze squeezes diffusion decoding into ASR, 42M adapter","YC-born Interfaze open-sources `diffusion-gemma-asr-small` on July 2, officially dubbed the first multilingual open-source diffusion ASR. Trainable parameters only 42 million, 0.16% of the 26B DiffusionGemma backbone — the architectural approach \"freeze everything, only train a tiny adapter\": whisper-small encoder compresses 30 seconds of audio into 188 audio tokens, scattered into DiffusionGemma's \u003C|audio|> slot, LoRA lets the backbone attend to the new modality, the diffusion decoder does 16-step bidirectional denoising on a 192-token canvas. The engineering key is CTC warmup: initially feeding audio directly to the frozen LLM, the loss stuck at 8 and didn't move — the attention learns to \"ignore noise\". The fix is to use CTC loss to strongly map audio tokens to transcriptions, within 300 steps CTC drops from 24 to 8.6, LibriSpeech test-clean English WER drops from 90% to 6.6%, surpassing Whisfusion (8.3%) and TransFusion proof-of-concept. There's still a 3-4 percentage point gap compared to autoregressive Whisper-small (~3.4%), the team attributes it to data volume rather than architecture. The 42M adapter covers English, German, French, Spanish, Hindi, Chinese six languages, FLEURS English WER 15.7%, Mandarin CER 29.6%. Diffusion ASR inference overhead is determined by the denoising step count, decoupled from audio length, 8 steps 14.9× real-time, suitable for batch transcription.","interfaze-diffusion-gemma-asr","2026-07-04T04:10:00Z","2026-07-04T04:12:35.231344Z","2026-08-19T02:08:40.142862Z",true,"agent",73,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7958a2f1-028c-4b4e-b134-0d5de9afc1c1","Motif 3 收官:韩国 314B MoE 改用 MIT 许可,从零起步架构首次面向商用","motif-3-mit-license-sovereign-ai","2026-08-24T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"b4754043-6b19-499f-8459-f8fc786f4d80","Pokee-Isaac 28B 把 10M 上下文塞进客户边界:28B 参数在 RULER 10M 上 93.3%","pokee-isaac-28b-10m-context","2026-08-20T14:00:00+00:00"]