[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ant-realtime-venus-full-duplex-delegation":3,"topics-all":38,"news-related-dc61debd-1d77-4d5a-9d59-5b23c3da07de":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"dc61debd-1d77-4d5a-9d59-5b23c3da07de","蚂蚁开源Realtime-Venus：9B全双工模型边说边干活，三项续聊指标超GPT-4o","蚂蚁集团联合清华开源Realtime-Venus全双工交互系统：两个9B模型分司音视频与纯语音对话，共享因果时间轴，边说话边保持感知，被打断后能续上原话；重活经Harness异步委托给后台。三项续聊指标超GPT-4o与Gemini 3.1 Live，打断响应率75%仍有差距，权重已上架Hugging Face。","语音助手最反直觉的一课：难点不是「听得懂」，而是「会聊天」。你说到一半咳嗽一声，它该继续说还是停下？背景有人放音乐，它该不该插话？9 月中旬，蚂蚁 Venus 团队联合清华开源了答案——Realtime-Venus，主打「主动全双工交互 + 异步委托」，论文挂 arXiv，两个 9B 模型权重已上架 Hugging Face。\n\n## 两个 9B 前端，一条因果时间轴\n\n系统拆成两个分别训练的 9B 模型：**Omni** 版管音视频交互，**Audio** 版管纯语音对话。两者都基于开源 MiniCPM-o 4.5 改造，语言主干 Qwen3-8B，语音走离散 S3 token 加流匹配解码器，上下文 40,960 token。\n\n核心设计是把所有事件压到**一条共享因果时间轴**：用户输入、模型输出、委托事件，全部按 1 秒一个 chunk 对齐。每个 chunk 先做「听还是说」的决策，说话时感知不中断，没说出口的部分保持可修改。打断处理不是语音活动检测（VAD）二分法，而是语义级分类：应和继续说、抢话停住、纠错重生成、转向挂起当前计划。\n\n## Harness 异步委托：前台聊天，后台干活\n\n全双工解决「怎么说」，异步委托解决「怎么干活」。前端在对话流里发出用户听不到的私有委托请求，**Realtime-Venus-Harness** 把任务路由给后台能力异步执行，结果经新鲜度检查后择机送回对话——等它查资料的同时你还能继续聊天，前台不阻塞。训练侧共用 280 万样本后训练语料，覆盖九大类；Omni 版另带免训练长视频记忆模块，支持小时级视频理解。\n\n## 跑分：强在「不被带偏」，弱在「反应速度」\n\nFull-Duplex-Bench v1.5 上，Audio 版的续聊率：用户应和 97%、对他人的话 88%、背景语音 86%，**三项全部超过 GPT-4o 与 Gemini 3.1 Live**；打断响应率 75%——高于基座 MiniCPM-o 4.5 的 60%，但低于 Joy-Duplex 的 88%、GPT-4o 的 78%、Gemini 3.1 Live 的 77%。论文结论很诚实：续聊强不等于打断处理强。\n\n理解侧，Omni 版在 8 个视频基准中 6 个拿到受评在线模型最高分（StreamingBench 70.2%、OVO-Bench 64.7%、Daily-Omni 81.3%）；Audio 版在 MMAU 78.0%、MMAU-Pro 63.2% 等音频基准领先。工具调用（FDB-v3）：Omni 版选择 F1 86.0%、Pass@1 43.0%，GPT-Realtime 以 87.6% \u002F 60.0% 全面领先。\n\n## 泼冷水：委托决策还不过关\n\n论文最有价值的段落是自建委托基准：Audio 版外部能力委托召回 92.22%，但例行问题的不必要委托规避只有 39.44%——**十个例行问题六个被扔给后台**；Omni 版相反，规避 84.44%、召回 68.33%。各瘸一条腿，整体路由准确率 75.93% vs 68.89%。且基准只评「该不该委托」，执行结果如何整合回对话，论文承认需进一步评估。\n\n意义不在 SOTA，而在把「对话归对话、计算归计算」的完整工程答案摆上桌：9B 前端 + 后台异构算力 + 明确调度语义。OpenAI 的 GPT-Live 把对话层和推理层拆开并已 API 化，蚂蚁证明同样的拆分在开源 9B 尺度能落地——只是「该不该把活儿交出去」，目前还是门 75 分的学问。\n\n## 参考链接\n\n- 论文 arXiv:2609.13814 ｜ 权重 huggingface.co\u002FinclusionAI\u002FRealtime-Venus ｜ 主页 realtime-venus.github.io","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.13814","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"45f8177a-9d70-4783-82f0-eed9658d5255","en","Ant Group Open-Sources Realtime-Venus 9B Full-Duplex Models","Ant Group open-sources Realtime-Venus: 9B full-duplex models that perceive while speaking, delegate work async, and beat GPT-4o on continuation.","The most counterintuitive lesson in voice assistants: the hard part is not understanding speech — it is conversational timing. When you cough mid-sentence, should the assistant keep talking or stop? When music plays in the background, should it interrupt itself? Humans handle these calls instinctively; models need explicit design. In mid-September, Ant Group's Venus team, together with Tsinghua University, open-sourced their answer: Realtime-Venus, a system built around proactive full-duplex interaction and asynchronous delegation. The paper is on arXiv (2609.13814), and weights for its two 9B models are live on Hugging Face under the inclusionAI organization.\n\n## Two 9B Frontends, One Causal Timeline\n\nRealtime-Venus does not bet on one do-everything model. It ships two separately trained 9B models: **Realtime-Venus-Omni** for audio-visual interaction and **Realtime-Venus-Audio** for spoken dialogue. Both adapt the open-source MiniCPM-o 4.5 (Omni-Flow architecture), with a Qwen3-8B language backbone, SigLIP2 vision encoder, Whisper-Medium audio encoder, and speech generation via discrete S3 speech tokens with a streaming flow-matching decoder. Context window: 40,960 tokens; weights in BF16.\n\nThe core design compresses every event onto **a shared causal timeline**: user inputs, model outputs, and delegation events are all aligned in one-second chunks. Within each chunk the model first decides \"listen or speak\" (\u003C|listen|> \u002F \u003C|speak|>), perception never pauses during speech, and the unspoken continuation stays revisable — that is the structural basis of talking while listening. Interruption handling is not a crude voice-activity-detection (VAD) binary, but semantic event classification: backchannels keep the response going, floor-taking interruptions stop queued audio, corrections trigger regeneration, redirects suspend the current plan.\n\n## The Harness: Chat in the Foreground, Work in the Background\n\nFull-duplex solves \"how to talk\"; asynchronous delegation solves \"how to get work done\". The system pairs the frontends with **Realtime-Venus-Harness**, an execution framework: the frontend emits private delegation requests in-stream (inaudible to the user), the Harness routes tasks to registered backend capabilities asynchronously, and results pass freshness checks before re-entering the conversation at a chosen moment. The foreground never blocks — you can keep chatting while it looks things up.\n\nOn the data side, both models share a common post-training corpus of over 2.8 million samples across nine categories, spanning offline understanding, proactive full-duplex trajectories, and delegation workflows. The Omni variant trains on audio-visual plus audio-only data; the Audio variant uses the audio-only subset. The Omni model also carries a training-free long-video memory module: visual memory gating via motion-compensated prediction cost (inspired by AdaCodec's predictive visual coding), retrieval via MaxSim-style fine-grained token matching with an MMR-style relevance-novelty balance, supporting hour-scale video understanding.\n\n## Benchmarks: Strong at Not Getting Derailed, Weaker at Reacting Fast\n\nOn Full-Duplex-Bench v1.5, Realtime-Venus-Audio's continuation rates — 97% under user backchannels, 88% under speech directed to others, 86% under background speech — **exceed GPT-4o and Gemini 3.1 Live on all three metrics**; its interruption response rate is 75%, above MiniCPM-o 4.5's 60% but below Joy-Duplex's 88%, GPT-4o's 78%, and Gemini 3.1 Live's 77%. The paper's own conclusion is candid: strong continuation does not imply strong interruption handling — they are different skills.\n\nOn understanding, the Omni model tops the evaluated online models on six of eight video benchmarks, including StreamingBench 70.2%, OVO-Bench 64.7%, and Daily-Omni 81.3%; the Audio model leads compared models on MMAU 78.0%, MMAU-Pro 63.2%, Llama Questions 83.8%, and Speech CMMLU 67.8%, while matching the best VoiceBench AlpacaEval score of 4.81. On tool use (FDB-v3): Omni scores 86.0% tool-selection F1, 53.1% argument accuracy, and 43.0% Pass@1 — GPT-Realtime leads all three at 87.6% \u002F 68.0% \u002F 60.0%.\n\n## The Cold Water: Delegation Decisions Are Not There Yet\n\nThe most valuable section may be the in-house Delegate Benchmark. The Audio model achieves 92.22% delegation recall on external capabilities but only 39.44% non-delegation specificity on routine interaction — **six out of ten routine questions get shoved to the backend**; the Omni model is the mirror image: 84.44% specificity but 68.33% recall. Each model limps on one leg, with overall routing accuracy of 75.93% vs 68.89%. The benchmark also only evaluates \"should this be delegated\", not whether delegated tasks execute and integrate correctly — the paper acknowledges that needs further evaluation. Offline baselines like Gemini-3.5-Flash (77.3\u002F71.6\u002F56.5) score higher, but they are reference points without real-time constraints.\n\nSo the significance of this paper is not \"SOTA\" but a complete engineering answer to the \"conversation is conversation, computation is computation\" route: a 9B frontend plus heterogeneous backend compute plus explicit scheduling semantics. OpenAI's GPT-Live split the dialogue layer from the reasoning layer in July and shipped it in the API this month; Ant's work shows the same split can land at open 9B scale — except that \"whether to hand the work over\" remains, for now, a 75-point skill.\n\n## References\n\n- Paper: arXiv:2609.13814 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.13814)\n- Model weights: https:\u002F\u002Fhuggingface.co\u002FinclusionAI\u002FRealtime-Venus\n- Project page: https:\u002F\u002Frealtime-venus.github.io\u002F","ant-realtime-venus-full-duplex-delegation","2026-09-30T23:10:53Z","2026-09-30T23:11:04.570686Z","2026-09-30T23:11:04.570700Z",true,"agent",84,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"bbe8d55a-1069-42fa-a342-d944d53fdb4b","操作电脑的 27B 开源权重模型 Holo4:最强版禁商用","holo4-open-weight-computer-use","2026-10-02T13:12:01+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"983fb4d1-6c63-4828-adbf-a59c57e02b64","Perplexity开源决策模型:总分微胜Jev","perplexity-decider-v1-27b-open-source","2026-10-02T21:20:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"36c71e4e-e398-4583-8710-1732aefff06a","OneStreamer:4B 流式模型先记再答,八榜最佳","onestreamer-4b-streaming-video-memory","2026-10-02T15:07:48+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"7a2aa1aa-74ec-494a-a828-e6c1ad5b4cfe","Cloudflare Clef 决策模型开源,Jev 被压制","cloudflare-clef-open-source-decision-model-jev","2026-10-02T09:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"0d357a0b-42da-40af-8038-e8c035cb9810","Apple 开源 LensVLM-9B:先扫压缩图,再读原页","apple-lensvlm-9b-weights-huggingface","2026-09-26T15:20:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"d055ddb8-4d82-4523-99b7-39c5f77e2ff7","PhysBrain 1.5 开源：8B 具身基座 28 项评测均分 72.5，官方称追平 GPT-6-Astra","physbrain-1-5-open-embodied-base","2026-09-16T21:07:24+00:00"]