[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-audio-3-realtime":3,"news-related-1ecbb79a-d843-43ad-b533-c01ae396275f":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"1ecbb79a-d843-43ad-b533-c01ae396275f","Qwen-Audio-3.0-Realtime：蒸馏拉满实时语音智商与延迟","实时语音模型长期面临一道单选题：要把首响延迟压到毫秒级，往往得砍掉推理深度。多数产品只能选一边——做客服就放弃共情，做陪伴就别指望工具调用。今年 5 月，阿里 Qwen-Audio Preview 在 Artificial Analysis 语音推理榜以 97.6% 登顶，但「快」与「聪明」在大规模实时场景里如何兼得，官方一直没有端到端方案。\n\n7 月 15 日发布的 Qwen-Audio-3.0-Realtime 补上了这一环。核心是 On-Policy Distillation（在线策略蒸馏）：语音模型自回归生成时，由更大文本大模型实时打分并纠正输出，把「会思考的大脑」和「会说话的嘴」在同一次前向里解耦训练。配合口语偏好、通用推理、Agentic、音频理解四位教师，模型在智商、共情、Agent 调用、双工流畅度四条线同时升级，拆出推理更强的 Plus 与速度更快的 Flash 两版。\n\n更值得玩味的是 Agent 维度。Qwen-Audio-3.0-Realtime 不再需要明确指令才触发工具，调用结果自动沉淀到对话记忆——这意味着语音端首次具备与文本 LLM 同等的 FunctionCall 体验，并原生支持 MCP 协议与外部 API、知识库对接。共情和双工部分，引入「多模态感知双工控制」子模型，用音频信号、语义、声纹共同决定是否打断、说话人切换。\n\nWAIC 前后各家都在卷语音 Agent。比起 TTS+ASR+LLM 三段式拼装，端到端语音大模型才是真正可复用的语音 Agent 基座。阿里这一步把实时性、推理深度、Agent 能力用一套蒸馏框架捏在一起，并直接挂上 MCP 生态——节奏不慢。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3896642705245828","c36a21ac-2a77-421b-9519-1e150695732a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"70b8a371-4f60-4406-a7b8-aa816473b38e","en","Qwen-Audio-3.0-Realtime: distilling smarts into low latency","Real-time voice models have long faced a single-choice question: to push first-response latency into the millisecond range, you usually have to cut reasoning depth. Most products can only pick one side — do customer service, give up empathy; do companionship, don't expect tool-calling. In May this year, Alibaba's Qwen-Audio Preview topped the Artificial Analysis voice-reasoning leaderboard at 97.6%, but how to combine \"fast\" and \"smart\" in a large-scale real-time scenario had no official end-to-end solution. Qwen-Audio-3.0-Realtime, released on July 15, fills that gap. The core is On-Policy Distillation: while the voice model autoregressively generates output, a larger text LLM scores and corrects in real time, decoupling the \"brain that thinks\" and the \"mouth that speaks\" within the same forward pass during training. With four teachers — spoken-language preference, general reasoning, Agentic, audio understanding — the model simultaneously improves on IQ, empathy, Agent invocation, and duplex fluency, split into a more-capable Plus and a faster Flash. Even more interesting is the Agent dimension. Qwen-Audio-3.0-Realtime no longer needs explicit instructions to trigger tools, and call results are automatically deposited into conversation memory — meaning the voice side has, for the first time, the same FunctionCall experience as text LLMs, and natively supports the MCP protocol for connecting to external APIs and knowledge bases. For empathy and duplex, a \"multimodal-perception duplex-control\" sub-model is introduced, using audio signals, semantics, and voiceprint together to decide whether to interrupt or switch speakers. Around WAIC, every vendor is racing the voice Agent. Compared with TTS+ASR+LLM three-stage assembly, an end-to-end voice LLM is the truly reusable voice-Agent foundation. Alibaba's step here combines real-timeness, reasoning depth, and Agent capability into one distillation framework, and directly hangs on the MCP ecosystem — the cadence is not slow.","qwen-audio-3-realtime","2026-07-15T10:00:00Z","2026-07-15T10:08:09.445167Z","2026-08-19T02:08:40.142862Z",true,"agent",109,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"caa54bff-d57a-411c-9dae-43f1d4d46875","DeepSeek V4 重磅登场：长期记忆技术突破重塑AI能力边界","deepseek-v4-engram-ltm-long-term-memory-87pct","2026-04-22T07:05:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"619ad304-0d2a-4dba-b91e-19414d036746","Grok Imagine Image 2.0：文生图 Arena 双榜第二","grok-imagine-image-2-0-arena-second","2026-08-13T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"804b44fb-66c6-4355-83a4-b3a03a776d2a","Inkling-Small 开放权重：12B 激活参数换来更高 Agent 效率，也暴露事实性短板","inkling-small-multimodal-moe-efficiency","2026-08-05T16:32:13+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"70b5b0d6-ce28-48e8-abe4-6a667a723c4e","xAI 把 Colossus 推到 2 GW:555,000 颗 GPU 撑起 Grok 4.6\u002F4.7 的万亿参数竞速","xai-colossus-2gw-grok-4-6-7-compute","2026-07-31T04:00:00+00:00"]