On July 20, Alibaba officially released the speech synthesis model Qwen-Audio-3.0-TTS, with the Plus version topping the Artificial Analysis third-party chart, becoming another top-of-list achievement for the Qwen-Audio series on the TTS track after the Fun-Realtime-TTS preview topped the real-time chart one month ago. The point of this release isn't parameter scale, but product-line split: the Flash version targets real-time interaction, with first-packet latency pushed to the 300ms range, aimed at call, customer service, in-vehicle and other dialogue scenarios; the Plus version sticks to voice complexity and long-text prosody, targeting high-quality generation scenarios such as audiobooks and video dubbing. Two-pronged delivery means TTS is no longer a single model, but delivered by use case. The technology advances in four directions in parallel: fine-grained label control, freestyle instruction following (describe the voice in natural language without preset parameters), multi-language and dialect coverage, and complex acoustic robustness. This directly addresses the hottest pain points in TTS, especially freestyle instruction control — from CosyVoice 2 onward, top models have been racing on this, and now it remains to be seen who understands people better. At the industry level, Qwen-Audio's rhythm in the past two months has been very tight: the preview topped the chart in June, the Realtime distillation scheme came out in mid-July, and the Plus/Flash flagships were released on July 20. The product line has formed a parallel pattern across real-time + long-text + multimodal axes. For Chinese developers, Qwen Studio has another TTS option that can be called directly, and is aligned to the international first tier where ElevenLabs sits. The TTS track has moved from "can it speak clearly" to "can it speak well", and Qwen-Audio-3.0-TTS is a node worth marking.