Qwen quietly pushed the all-modality price war to a new level on September 18. Qwen3.8-Omni-Flash shipped with a 1M-token context for text, image, audio, and video inputs; the audio API rate drops 98% and the combined audio-video input rate drops 93%. The move squeezes the mid-tier survival space of Gemini 3.8 Flash and other Flash-tier rivals.

What Actually Changed in This Upgrade

The official benchmark set shows Qwen3.8-Omni-Flash averaging over 26% above Qwen3.5-Omni-Plus across 30 evaluations. The biggest gains sit in the audio-video Agent category: WildClawBench-MM lifts 36.5 points, AgenticVBench lifts 22.3 points, and UniClawBench scores 69.6.

In meeting scenarios, AliMeeting DER/cpWER drops from 88.11/89.61 to 3.35/17.18, a near-order-of-magnitude improvement. The official statement now reads "audio-video capability close to Gemini 3.8 Flash, audio capability overall exceeds Gemini 3.8 Flash", a direct confrontation rather than the prior off-angle comparisons against Opus 4.8.

The more consequential Agentic long-video numbers: OmniVideoBench accuracy lifts from 63.4 to 67.8, while token consumption on the same task drops from 145,736 to 79,117, a 45.7% reduction. Capability up 7% with tokens nearly halved, the marginal-return curve is what makes the new price possible.

Why the Price Could Drop This Hard

All-modality models never really fell in price before because tokens explode: one hour of meeting video at 1fps plus audio transcription can produce millions of tokens. Qwen3.8-Omni-Flash solves this in two layers. First, internal attention and caching optimizations cut Agentic long-video tokens by 45%. Second, the team simultaneously open-sourced two companion pieces, Qwen-MM-Plugins (on-demand perception, tool use, and execution for long workflows) and Qwen-Live Harness (sustained real-time all-modality interaction), letting enterprises run Qwen3.8-Omni-Flash on local or private cloud rather than paying public-API bills. For organizations consuming tens of millions of tokens per month this is decisive.

Realtime: The First Open-Source All-Modality Model with Sound Localization

Shipped alongside the base model, Qwen3.8-Omni-Flash-Realtime focuses on stream-in real-time response and is described by the team as the first open all-modality model with sound localization, fusing spatial audio and visual information to estimate sound-source direction and distance. This is a real need in robotics, ADAS, and AR devices where users will not constrain themselves to voice-only commands in noisy settings.

Realtime also supports real-time language tutoring, jointly modeling pronunciation and semantics to understand accented and non-standard expressions. Previously this had to be stitched together from ElevenLabs-class TTS plus a third-party LLM; Qwen has collapsed the two into a single model, a notable competitive signal for consumer agent products like AI glasses and AI toys.

Self-Optimization: Another Agent + RL Sample

In an officially disclosed optimization experiment, Qwen3.8-Omni-Flash tried to improve Qwen2.5-Omni-3B Sichuan-dialect recognition in 12 hours: it autonomously selected an evaluation set, built 3413 training samples over 4 rounds, and dropped character error rate from 25.79% to 15.30% (a 40.7% relative drop).

This is not self-training in the strict sense but an Agent + RL data-synthesis-plus-optimization sample: the model uses tools to generate data, then trains a sibling model on that data. For companies fine-tuning small models, this is a direct prompt: the bottleneck for your 1B-3B model may no longer be "do we have enough data" but "can an Agent produce high-quality synthetic data in our domain".

Industry Impact

All-modality just dropped to the 1% price band, the same tempo as GPT-5.6 and DeepSeek V4-Flash. But all-modality burns more than text or vision alone, and the Flash-tier breakeven line was thought hard to push. In the short run, Gemini 3.8 Flash and Claude Sonnet 5.5 will have to answer with their own pricing moves. In the medium run, this band moves all-modality from "demo grade" to "product grade" projects that were once blocked by compute and API costs in sub-200-dollar BOM AI glasses can be re-budgeted.

So What

Teams working on AI glasses, ADAS, robotics, or all-modality Agents should pull Qwen3.8-Omni-Flash plus Qwen-MM-Plugins plus Qwen-Live Harness onto a new branch and re-evaluate end to end. It is not just a capability bump; the unit price has crossed into the range that can be assumed into a product configuration. Flash-tier competitors will need a pricing response in the next few weeks rather than waiting for the next benchmark.