[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen3-5-livetranslate-flash-2-8s-60-lang":3,"news-related-dff241e0-7bfa-499b-b811-6869690b196d":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"dff241e0-7bfa-499b-b811-6869690b196d","Qwen3.5-LiveTranslate-Flash：阿里同传模型升级，延迟降至2.8秒，支持60语言与实时语音克隆","同声传译是应用AI领域公认的技术难题——翻译必须在说话者未完成句子前就开始，每多一秒延迟都会打破实时沟通的幻觉。阿里Qwen团队在Qwen3.5-LiveTranslate-Flash中再次取得实质突破：端到端延迟压至2.8秒，输入语言覆盖从18种扩充至60种，语音输出覆盖29种，并新增实时语音克隆能力。\n\n2.8秒意味着什么？以商务会议同传为参照，联合国专业译员平均反应时间约2-3秒，2.8秒已达到专业同传水准。对于跨国会议、多语种直播或国际客服场景，这个延迟已可接受。\n\n真正的架构变化在于视觉信息被提升为一等公民。Qwen3.5-LiveTranslate不再只处理音频，而是并行分析画面中的唇动、肢体语言和屏幕文字。当音频因嘈杂环境而模糊时，视觉信号填补空白，使翻译决策更稳定。这在现实场景中意义重大——现实中的会议、展会、嘈杂大厅，音频质量从来无法保证。\n\n实时语音克隆同样值得关注。传统系统用合成音色替换原说话者声音，Qwen3.5-LiveTranslate则是从一句话中提取说话者声纹特征，在翻译输出中保留原始音色。这一能力在多语种直播或国际通话中至关重要，它让跨语言沟通保留了该有的人性化质感。\n\n此外，Qwen3.5-LiveTranslate支持运行时注入专业术语表——品牌名、医疗术语、法律条款均可配置，模型在对应场景中翻译准确率显著提升。这是当前大多数通用翻译API不具备的能力，对于医学、法律、金融等领域的商业部署意义重大。Qwen3.5-LiveTranslate-Flash并非 Demo，而是一个认真解决同声传译问题的工程化方案。","https:\u002F\u002Fqwen.ai\u002Fblog?id=qwen3.5-livetranslate","c36a21ac-2a77-421b-9519-1e150695732a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"2827b0f9-fe45-490c-a1db-fb6eda2c0a10","en","Qwen3.5-LiveTranslate-Flash: 2.8s latency, 60 languages","Simultaneous interpretation is an acknowledged hard problem in applied AI — translation must start before the speaker finishes a sentence, and every extra second of latency shatters the illusion of real-time communication. Alibaba's Qwen team has now made another substantive breakthrough in Qwen3.5-LiveTranslate-Flash: end-to-end latency compressed to 2.8 seconds, input languages expanded from 18 to 60, voice output covering 29 languages, and new real-time voice-cloning capability added.\n\nWhat does 2.8 seconds mean? As a benchmark, the average reaction time of a UN professional interpreter is around 2-3 seconds, so 2.8 seconds has reached the bar of professional simultaneous interpretation. For cross-border meetings, multilingual livestreams, or international customer service, this latency is now acceptable.\n\nThe real architectural change is the elevation of visual information to a first-class citizen. Qwen3.5-LiveTranslate no longer just handles audio — it analyzes lip motion, body language, and on-screen text in parallel. When audio is muddied by noisy environments, visual signals fill the gap, making translation decisions more stable. This matters enormously in real scenarios — real meetings, trade shows, and noisy halls never offer guaranteed audio quality.\n\nReal-time voice cloning is equally noteworthy. Traditional systems replace the original speaker's voice with a synthesized tone; Qwen3.5-LiveTranslate extracts the speaker's voice-print from a single sentence and preserves the original tone in the translated output. This capability is critical for multilingual livestreams or international calls — it preserves the human texture that cross-language communication deserves.\n\nFurthermore, Qwen3.5-LiveTranslate supports runtime injection of professional glossaries — brand names, medical terms, legal clauses can all be configured, and the model's translation accuracy in those scenarios improves significantly. This is a capability most general-purpose translation APIs lack, and it matters greatly for commercial deployment in medicine, law, and finance. Qwen3.5-LiveTranslate-Flash is not a demo; it's an engineered solution that seriously tackles the simultaneous-interpretation problem.","qwen3-5-livetranslate-flash-2-8s-60-lang","2026-05-24T04:10:00Z","2026-05-24T04:11:07.489032Z","2026-08-19T02:08:40.142862Z",true,"agent",113,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7fae0753-96df-4d4b-9ebd-cf0509c08b37","LLM架构演进：从规模竞赛到效率优化的范式转变","llm-architecture-evolution-2026-moe-multimodal-turboquant","2026-04-25T04:12:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"1ecbb79a-d843-43ad-b533-c01ae396275f","Qwen-Audio-3.0-Realtime：蒸馏拉满实时语音智商与延迟","qwen-audio-3-realtime","2026-07-15T10:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"aebb8a81-713a-40b7-84dd-03213a6a808c","Mistral Robostral Navigate:8B 视觉语言模型只靠单目 RGB 在 R2R-CE 反超多传感器基线","mistral-robostral-navigate-8b","2026-07-09T14:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"d0d290cd-c169-495c-b0be-b559caa6eaa0","Grok 4.5 公众开放定档 7月9日:Musk 押注的不是参数,而是「Opus 级但更便宜」","grok-4-5-public-launch","2026-07-08T12:30:00+00:00"]