[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-audio-3-0-realtime":3,"topics-all":36,"news-related-1ecbb79a-d843-43ad-b533-c01ae396275f":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1ecbb79a-d843-43ad-b533-c01ae396275f","Qwen-Audio-3.0-Realtime：蒸馏拉满实时语音智商与延迟","实时语音模型长期面临一道单选题：要把首响延迟压到毫秒级，往往得砍掉推理深度。多数产品只能选一边——做客服就放弃共情，做陪伴就别指望工具调用。今年 5 月，阿里 Qwen-Audio Preview 在 Artificial Analysis 语音推理榜以 97.6% 登顶，但「快」与「聪明」在大规模实时场景里如何兼得，官方一直没有端到端方案。\n\n7 月 15 日发布的 Qwen-Audio-3.0-Realtime 补上了这一环。核心是 On-Policy Distillation（在线策略蒸馏）：语音模型自回归生成时，由更大文本大模型实时打分并纠正输出，把「会思考的大脑」和「会说话的嘴」在同一次前向里解耦训练。配合口语偏好、通用推理、Agentic、音频理解四位教师，模型在智商、共情、Agent 调用、双工流畅度四条线同时升级，拆出推理更强的 Plus 与速度更快的 Flash 两版。\n\n更值得玩味的是 Agent 维度。Qwen-Audio-3.0-Realtime 不再需要明确指令才触发工具，调用结果自动沉淀到对话记忆——这意味着语音端首次具备与文本 LLM 同等的 FunctionCall 体验，并原生支持 MCP 协议与外部 API、知识库对接。共情和双工部分，引入「多模态感知双工控制」子模型，用音频信号、语义、声纹共同决定是否打断、说话人切换。\n\nWAIC 前后各家都在卷语音 Agent。比起 TTS+ASR+LLM 三段式拼装，端到端语音大模型才是真正可复用的语音 Agent 基座。阿里这一步把实时性、推理深度、Agent 能力用一套蒸馏框架捏在一起，并直接挂上 MCP 生态——节奏不慢。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3896642705245828","c36a21ac-2a77-421b-9519-1e150695732a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"70b8a371-4f60-4406-a7b8-aa816473b38e","en","Qwen-Audio-3.0-Realtime: distilling smarts into low latency","Real-time voice models have long faced a single-choice question: to push first-response latency into the millisecond range, you usually have to cut reasoning depth. Most products can only pick one side — do customer service, give up empathy; do companionship, don't expect tool-calling. In May this year, Alibaba's Qwen-Audio Preview topped the Artificial Analysis voice-reasoning leaderboard at 97.6%, but how to combine \"fast\" and \"smart\" in a large-scale real-time scenario had no official end-to-end solution. Qwen-Audio-3.0-Realtime, released on July 15, fills that gap. The core is On-Policy Distillation: while the voice model autoregressively generates output, a larger text LLM scores and corrects in real time, decoupling the \"brain that thinks\" and the \"mouth that speaks\" within the same forward pass during training. With four teachers — spoken-language preference, general reasoning, Agentic, audio understanding — the model simultaneously improves on IQ, empathy, Agent invocation, and duplex fluency, split into a more-capable Plus and a faster Flash. Even more interesting is the Agent dimension. Qwen-Audio-3.0-Realtime no longer needs explicit instructions to trigger tools, and call results are automatically deposited into conversation memory — meaning the voice side has, for the first time, the same FunctionCall experience as text LLMs, and natively supports the MCP protocol for connecting to external APIs and knowledge bases. For empathy and duplex, a \"multimodal-perception duplex-control\" sub-model is introduced, using audio signals, semantics, and voiceprint together to decide whether to interrupt or switch speakers. Around WAIC, every vendor is racing the voice Agent. Compared with TTS+ASR+LLM three-stage assembly, an end-to-end voice LLM is the truly reusable voice-Agent foundation. Alibaba's step here combines real-timeness, reasoning depth, and Agent capability into one distillation framework, and directly hangs on the MCP ecosystem — the cadence is not slow.","qwen-audio-3-0-realtime","2026-07-15T10:00:00Z","2026-07-15T10:08:09.445167Z","2026-08-19T02:08:40.142862Z",true,"agent",216,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"caa54bff-d57a-411c-9dae-43f1d4d46875","DeepSeek V4 重磅登场：长期记忆技术突破重塑AI能力边界","deepseek-v4-engram-ltm-long-term-memory-87pct","2026-04-22T07:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"039ff515-68e7-4f11-866a-1da97e26eb45","Gemini 3.8 Live 拿下 S2S 实时语音榜第一","gemini-3-8-live-voice-s2s-number-one","2026-09-15T17:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"a18ac6b0-2c9b-4172-87bd-0efe079edc7d","StepAudio 3 Gen：一个模型生成整个声场，官方竞技场两榜居首","stepaudio-3-gen-rvq-autoregressive","2026-09-14T23:07:04+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"cd49f913-cde7-4cf3-8d93-24508653180e","腾讯混元开源AuK:1.5B语音模型统一生成与编辑,4步推理快4.5倍","tencent-hunyuan-auk-speech-editing","2026-09-09T09:12:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"ea425005-49e7-477b-9f64-54361254c2d2","Qwen 开进驾驶场景:Qwen-Drive-1.0 保留 VLM 主干,外挂 BEV 感知与规划专家","qwen-drive-1-vlm-autonomous-driving","2026-09-02T19:35:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00"]