Alibaba used its 2026 Apsara Conference on Monday to redraw the LLM scaling map. The current flagship Qwen3.8-Max already runs at 2.4 trillion parameters; the newly announced Qwen4 is now in training, while Qwen4.5 and Qwen5 are planned to push the parameter count into the 5–10 trillion range, aimed directly at artificial superintelligence (ASI). The team put the "parameter scaling is still the core path to AGI" claim back on the table at a moment when Scaling Law has been repeatedly questioned, effectively a counter-vote.
What deserves more attention than raw scale is the recursive self-improvement (RSI) Qwen3.8-Max has demonstrated. With zero human involvement, the system iterated 33 rounds in a row — building its own training pipelines, generating training data, designing experiments and fixing defects — and lifted its Artificial Analysis agent benchmark score by 12.5%. The significance is not the number itself, but that "a model editing its own training pipeline" now has industry-grade evidence. Combined with the Pingtouge Zhenwu V900 chip — three times the performance of the previous generation, scalable to clusters of 500,000 cards, in volume from Q1 2027 — RSI is no longer an isolated algorithmic demo but the concrete glue for the model–chip–datacenter three-layer flywheel.
The multimodal lineup revealed at Apsara is denser than the LLM roadmap itself. Qwen3.8-Omni-Flash unifies video, audio, image and text; Wan3.0 pushes single-shot video generation to 30 seconds and accepts structured instructions; Qwen-Audio-3.1 covers ASR, TTS and real-time translation with per-character real-time translation latency compressed to 2.3 seconds; Qwen-Image-3.1, HappyOyster-2.0-Preview and other specialised models are advancing in parallel. A next-generation video generation model is scheduled for release in November — the first time Qwen has put a clear milestone on video, and its entry pass against the Veo/Sora tier.
Open source and commercialisation accelerating at the same time
Within a month of launch, the Qwen3.8 series surpassed 56 million downloads with more than 1,900 derivative models. Alibaba has now open-sourced over 460 Qwen models, with cumulative downloads above 3 billion and more than 300,000 derivative models. On the commercial side, Qwen3.8-Flash cuts training cost by nearly 90% through attention-mechanism optimisation; Perplexity, Airbnb, Pinterest and Reuters have built agents and proprietary models on top of Qwen. In Q2 2026 Alibaba's AI cloud and compute-services revenue reached roughly 48.4 billion yuan, up 45% year-on-year, while capex hit 67.7 billion yuan, up 75% year-on-year. CEO Eddie Wu set the target: by 2032, the global datacenter footprint operated by Alibaba Cloud will exceed 20 GW.
So what
Parameter expansion, RSI, in-house silicon and the 20 GW datacenter target are no longer four separate items — they have been woven into a single line by the same announcement. Scaling Law has not been invalidated; it has been replaced by a three-in-one paradigm: model self-iteration + co-designed hardware + compute at utility scale. For domestic peers, Qwen4 entering training means the real starting line for the next flagship cycle has been pulled to the 5–10 trillion tier. The gap will no longer be decided by model size, but by who first turns RSI into an industrial pipeline and who first makes the chip–datacenter–model loop work end to end.