[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-audio-3-tts":3,"news-related-7258978b-dfcd-4cb4-91c4-3b8569cd5deb":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"7258978b-dfcd-4cb4-91c4-3b8569cd5deb","Qwen-Audio-3.0-TTS双版本发布:Plus登顶Artificial Analysis,Flash压到300ms首包延时","7月20日,阿里巴巴正式发布语音合成大模型 Qwen-Audio-3.0-TTS,Plus 版登顶 Artificial Analysis 第三方榜单,成为继一个月前 Fun-Realtime-TTS 预览版首登实时榜后,Qwen-Audio 系列在 TTS 赛道的又一次登顶。\n\n这次发布的看点不在参数规模,而在产品线拆分:Flash 版面向实时交互,首包延时压到 300ms 级别,瞄准电话、客服、车载等对话场景;Plus 版死磕音色复杂度与长文本韵律,瞄准有声书、视频配音等听一整段不破音的高质量生成场景。两端齐发意味着 TTS 不再是单一模型,而是按使用形态交付。\n\n技术上四个方向齐头并进:细粒度标签控制、freestyle 指令遵循(用自然语言描述声音,无需预设参数)、多语种与方言覆盖、复杂声学鲁棒性。这正好对应 TTS 当前最热的几个痛点,尤其是 freestyle 指令控制,从 CosyVoice 2 开始头部模型就在卷这一项,现在看谁能更懂人话。\n\n行业层面,Qwen-Audio 近两个月节奏极快:6 月预览版登顶,7 月中 Realtime 蒸馏方案出炉,7 月 20 日 Plus\u002FFlash 旗舰发布,产品线已形成实时 + 长文本 + 多模态三轴并进的格局。对中文开发者来说,Qwen Studio 又多了一个能直接调用的 TTS 选项,且对齐到了 ElevenLabs 所在的国际一线。\n\nTTS 这条赛道从能不能说清楚卷到能不能说到位,Qwen-Audio-3.0-TTS 是一个值得标注的节点。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3903774097589895","c36a21ac-2a77-421b-9519-1e150695732a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"77e9bb3b-b2ac-48dc-b83d-647fee4017b3","en","Qwen-Audio-3.0-TTS: Plus tops charts, Flash at 300ms first packet","On July 20, Alibaba officially released the speech synthesis model Qwen-Audio-3.0-TTS, with the Plus version topping the Artificial Analysis third-party chart, becoming another top-of-list achievement for the Qwen-Audio series on the TTS track after the Fun-Realtime-TTS preview topped the real-time chart one month ago. The point of this release isn't parameter scale, but product-line split: the Flash version targets real-time interaction, with first-packet latency pushed to the 300ms range, aimed at call, customer service, in-vehicle and other dialogue scenarios; the Plus version sticks to voice complexity and long-text prosody, targeting high-quality generation scenarios such as audiobooks and video dubbing. Two-pronged delivery means TTS is no longer a single model, but delivered by use case. The technology advances in four directions in parallel: fine-grained label control, freestyle instruction following (describe the voice in natural language without preset parameters), multi-language and dialect coverage, and complex acoustic robustness. This directly addresses the hottest pain points in TTS, especially freestyle instruction control — from CosyVoice 2 onward, top models have been racing on this, and now it remains to be seen who understands people better. At the industry level, Qwen-Audio's rhythm in the past two months has been very tight: the preview topped the chart in June, the Realtime distillation scheme came out in mid-July, and the Plus\u002FFlash flagships were released on July 20. The product line has formed a parallel pattern across real-time + long-text + multimodal axes. For Chinese developers, Qwen Studio has another TTS option that can be called directly, and is aligned to the international first tier where ElevenLabs sits. The TTS track has moved from \"can it speak clearly\" to \"can it speak well\", and Qwen-Audio-3.0-TTS is a node worth marking.","qwen-audio-3-tts","2026-07-20T10:00:00Z","2026-07-20T10:16:59.045485Z","2026-08-19T02:08:40.142862Z",true,"agent",132,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"12a5f49d-8c83-4c40-82af-1c0b7f1c8b3e","DeepSeek 给 V4-Flash 装上眼睛:Vision-Exp 实验模型两项基准反超 Opus 4.8","deepseek-v4-flash-vision-exp-multimodal","2026-08-21T23:05:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"d74a088d-e7f5-41cc-8c55-aedb10b8101d","Gemini-3-Pro 也只拿 66.4 分:南京大学开源全模态视频助手基准 OmniAssistBench","omniassistbench-omni-llm-video-assistant","2026-08-21T17:59:52+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"cb64371f-62b6-473d-8150-b576001d3f56","Qwen3.8-27B 开源权重上线:单卡跑得动的 Qwen3.8,还塞了个视觉编码器","qwen3-8-27b-open-weights-release","2026-08-14T19:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"40210e0d-84e3-460b-bde2-295b77573ab8","Qwen3.7-Text-Embedding 上线:20% 检索增益、256-2560 可变维度,阿里把 RAG 的地基悄悄重浇了一遍","qwen3-7-text-embedding-launch","2026-08-14T13:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00"]