[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-netease-youdao-r2t2-streaming-asr":3,"topics-all":38,"news-related-e7ef7085-3e55-4336-8383-0c1f1dbdfbdd":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"e7ef7085-3e55-4336-8383-0c1f1dbdfbdd","有道开源 R2T2：不改口流式 ASR 上线","有道开源 R2T2：一款 2B 参数的真流式转录模型，200-600ms 平均延迟，输出 token 一旦交付就不再修改。在英语、汉语八个公开数据集上接近离线基准水平，带 vLLM 与 WebSocket 服务。","## 实时转录终于不闪字幕了\n\n过去十年实时 ASR 一直背着个难听的包袱:模型先说一句话,然后立刻改口。直播字幕闪一个 \"你好\",等半秒,变成 \"你好,很高兴见到你\";下游 NLP 流水线被持续漂移的文本逼到重复做状态机;LLM 智能体每收到一段新稿就要把之前的推理重做一遍。网易有道新开源的 [Confucius4-R2T2](https:\u002F\u002Fhuggingface.co\u002Fnetease-youdao\u002FConfucius4-R2T2) 押的方向相反:流式模型理应一锤定音,签完就改不动。\n\nR2T2 短写是 \"Real Real-Time Transcription\",这个绕口令本身是介绍词——市面上大多数 \"实时 ASR\" 实际上是伪流式,输出会随着后续音频被改写。R2T2 建立在开源的 Qwen3-ASR 之上([网易有道官方 X 公告](https:\u002F\u002Fx.com\u002FNetEase_Global\u002Fstatus\u002F2100220446004429217)给出 1.7B 参数的开源口径),在训练数据侧引入稳定前缀数据、强制时序对齐数据,以及 token 级音频切分,模型只把\"有把握\"的部分送出来,一旦确定就锁死。解码块的长度可以从 80 ms 到 2 s 自由配,业务方按场景自己拿延迟换吞吐。\n\n这个赌注在基准测试里站得稳。官方自报的英语流式结果(AMI、Giga-clean、LS-clean、SPGI、TED-LIUM 等)在 160 ms 块长下,与离线版 Qwen3-ASR 基线的差距只有约 2 个百分点;汉语端(Wenet-net、Wenet-meeting、SPEECHIO)领先离线约 0.9 到 2.0 个 CER,仍然在做真正的流式推理。注意一个细节:这些数字的前提是 R2T2 在生成期间**不被允许**回头修改自己已经说出口的 token。官方把达成这一点的关键归到一个叫 Longest Stable Prefix(LSP)的学习范式——决定什么时候这段前缀已经稳得可以输出,什么时候需要再等等更多音频上下文。\n\n下游稳定是个比基准更值钱的胜利。直播字幕不再因为说话人继续开口而跳字;接 ASR 流水的 LLM 智能体,可以在\"已确认文本\"上挂工具调用逻辑,不用每 250 ms 重跑一次状态机;同声传译的对齐缓冲保持单调。客服、无障碍字幕、语音智能体这些最被实时转写折磨的场景,正好是 R2T2 最对准的目标。\n\n让它不只是一个演示,是把它做成开源可投产。两个抓手:代码以 Apache 2.0 释出,vLLM 推流后端再加一个现成的 WebSocket 服务(`ws_server.py`)和一个参考客户端;2B 参数权重(基座是 Qwen3-ASR-1.7B)走网易的研究友好许可证。Docker 路径也省心——用官方 `qwenllm\u002Fqwen3-asr` 镜像起好容器,`ws:\u002F\u002Flocalhost:8272\u002Fasr_stream_api_v1` 直接接 16 kHz 单声道 PCM,每帧约 160 ms。\n\n往大图里看,2026 年的流式 ASR 车道已经闹起来:微软 VibeVoice 7B 主攻说话人分离,Meta 的 Muse Voice Transcribe 把流式、20+ 说话人分离、端点检测塞进一个模型,NVIDIA Nemotron 3.5 ASR 把 40 种语言照顾到。R2T2 走了另一条互补赛道:多语种流式 + **append-only 稳定输出**,不在说话人和语种上堆。官方口径说自己\"在开源流式 ASR 中达到 SOTA,延迟上和闭源系统有竞争力\"。真正决定它分量的是接下来 6 个月的端到端数字:当下游的语音智能体不再是\"听说错再改一遍\"而是\"听一次就照办\",这条流水线到底能省下多少工程?这个数字才能决定\"稳定前缀\"会不会成为所有后续流式模型的默认能力。\n","https:\u002F\u002Fhuggingface.co\u002Fnetease-youdao\u002FConfucius4-R2T2","d5d9b417-5226-421d-a92e-e33e0ceb74f9",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"c4455801-0273-477d-9e92-00b4addfa439","asr",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"da116f20-377f-449f-9382-821d56f6fa64","streaming",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"0d37b28f-3e21-49fd-9ab9-6bcca2d2763c","en","Youdao's R2T2: streaming ASR that never rewrites past output","NetEase Youdao open-sourced R2T2: a 2B true-streaming ASR with append-only output. 200-600 ms latency, vLLM and WebSocket bundled, Apache 2.0.","## Why real-time transcription finally stops flickering\n\nFor a decade, real-time ASR has lived with an annoying tax: the model speaks, then corrects itself. A live subtitle flashes \"Hel-\", waits half a second, then rewrites to \"Hello there\". Downstream NLP pipelines choke on text that keeps shifting underneath them, and LLM agents re-plan on every new draft. NetEase Youdao's newly open-sourced [Confucius4-R2T2](https:\u002F\u002Fhuggingface.co\u002Fnetease-youdao\u002FConfucius4-R2T2) bets that streaming models should commit to a stable prefix and never rewrite.\n\nR2T2 stands for \"Real Real-Time Transcription\", a deliberate wordplay: most \"real-time\" ASR on the market is actually pseudo-streaming, where the model revises its previous output as more audio arrives. R2T2 is built on the open Qwen3-ASR base (per the [NetEase Youdao X announcement](https:\u002F\u002Fx.com\u002FNetEase_Global\u002Fstatus\u002F2100220446004429217)) and adds a training pipeline that includes stable-prefix data, forced time-alignment data, and token-level audio segmentation. The result is a model that exposes only the parts of the transcript it is confident about, and once committed, those tokens never move. Decoding chunks are configurable from 80 ms to 2 s, letting operators trade off latency against throughput per application.\n\nThe headline numbers from the model card land in the right place for that bet. Average streaming latency sits between 200 and 600 ms; on the English benchmark suite (AMI, Giga-clean, LS-clean, SPGI, TED-LIUM, etc.) R2T2 at 160 ms chunks is within ~2 percentage points of a fully offline Qwen3-ASR baseline despite never being allowed to revise its own output. On the Chinese benchmark set (Wenet-net, Wenet-meeting, SPEECHIO) it sits 0.9-2.0 CER points behind offline Qwen3-ASR while still running as a true streaming system, with the proprietary Commercial B ASR service only fractionally ahead. The model's own evaluation calls out a specific architecture choice that makes this work: a Longest Stable Prefix (LSP) learning paradigm that decides when a prefix is safe to emit versus when more audio context is needed.\n\nThe practical consequence is downstream stability. Live captioning no longer flickers when the speaker keeps talking. LLM agents consuming the stream can hang tool-calling logic off \"confirmed text\" without re-running their state machine every 250 ms. Simultaneous translation pipelines keep their alignment buffer monotonic. For real-time call-center, accessibility, and voice-agent use cases, this is exactly the property people have wanted since Whisper dropped streaming output and stopped being usable for live UX.\n\nTwo pieces of the design make this an open-weights story, not just a tech demo. The inference code is released under Apache 2.0 with a vLLM streaming backend and a WebSocket server (`ws_server.py`) that ships with a reference client. The 2B-parameter checkpoint (built on Qwen3-ASR-1.7B) is published under NetEase's research-friendly license. The release also bundles Docker integration with the official `qwenllm\u002Fqwen3-asr` image, so a developer can spin up `ws:\u002F\u002Flocalhost:8272\u002Fasr_stream_api_v1` and stream 16 kHz mono PCM in roughly 160 ms frames.\n\nFor the bigger picture, the streaming ASR lane has been getting noisy in 2026: Microsoft's VibeVoice 7B attacked speaker diarization; Meta's Muse Voice Transcribe folded streaming, 20+ speaker separation, and endpoint detection into one model; NVIDIA's Nemotron 3.5 ASR covered 40 languages. R2T2 occupies a complementary niche: multilingual streaming recognition with append-only stable output rather than diarization or language breadth. On its own benchmarks it claims SOTA among open-weight streaming ASR, competitive with closed-source systems on latency. The interesting comparison ahead is end-to-end: how much downstream rework do real voice-agent pipelines save when their transcript input never rewrites? That number is what will decide whether \"stable prefix\" becomes a default feature of every streaming model shipped after this.\n","netease-youdao-r2t2-streaming-asr","2026-09-21T09:01:41Z","2026-09-21T09:06:15.952095Z","2026-09-21T09:06:15.952104Z",true,"agent",23,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"cb5ee922-ad27-4a5a-9b2b-8382903876df","Mozilla 把模型选择权交还给用户:Mistral Small 4 进 Firefox 默认菜单","mistral-small-4-firefox-smart-window-beta","2026-09-22T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"d056f67b-7e0d-4e44-8d39-e31ea50deeae","Bonsai 2 27B 三元压缩:Qwen3.8 压到 5.9 GB,benchmark 留存 98.2%","bonsai-2-27b-ternary-qwen3-8-compression","2026-09-17T15:47:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"d055ddb8-4d82-4523-99b7-39c5f77e2ff7","PhysBrain 1.5 开源：8B 具身基座 28 项评测均分 72.5，官方称追平 GPT-6-Astra","physbrain-1-5-open-embodied-base","2026-09-16T21:07:24+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"21fe3c11-4ba4-4801-b6fc-60c4ae559dc1","Yandex 逆流开源:35B 参数的 T5 MoE,每个 token 只激活 0.6B","yandex-aliceai-t5-sparse-moe","2026-09-16T19:11:43+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"30fca629-bace-4832-9789-b44aa8c8989d","学生团队从零训出开源 7B 模型 ZGCM-1:数学推理硬刚 235B 前沿","zgcm-1-open-7b-foundation-model","2026-09-15T19:10:00+00:00"]