Mozilla published version 1.1 of the State of Open Source AI report on September 15, and the headline number is sharp: the capability gap between Chinese open-weight models and US frontier closed-source models has compressed to about 4.4 months.

Three points on the leaderboard, seven times the cost

Drawing on the September 1 cut of the Artificial Analysis Intelligence Index, Anthropic's Claude Fable 5 sits on top at 62 points, with Moonshot's Kimi K3 at 60. The score gap is just three points. The price gap is far wider. Fable 5 lists at $10 input / $50 output per million tokens; Kimi K3 lists at $3 / $15. Run the same workload through both, and Kimi K3's effective spend lands near 30% of Fable 5's. Z.ai's GLM-5.3 pushes the cost curve further down: $0.075 input per million tokens, the cheapest seat at this capability tier.

Mozilla CTO Raffi Krikorian frames that gap as a premium that lives almost entirely in the 8-to-12-hour professional task band. Below that, open weights do the job at a fraction of the cost.

Twelve hours is still the closed-source moat

Citing METR Time Horizon 1.1 (50% reliable-success criterion), the report lines both sides up on the same time axis: the best closed-source model now reliably completes a ~12-hour task; the best open-weight model reaches about 7 hours, a task-length ratio of roughly 1.7x. Krikorian's call is that four months from now, open weights catch 12 hours while closed-source pushes to ~20 hours, meaning the 8-to-12-hour band is exactly where closed-source still earns its premium today.

That call is already showing up in routing decisions. DoorDash has moved day-to-day workloads to Kimi while reserving Fable for more complex jobs. On OpenRouter's August token-volume ranking, eight of the top ten models are open-weight, and seven of the eight are built by Chinese teams.

August was the first month an open-weight model led OpenRouter

The most striking data point in the report sits in the OpenRouter section. On August 3, DeepSeek took over the weekly-request crown from Google, which had held it for 51 straight weeks. From 80 billion weekly requests in September 2025 to 1.01 trillion in August 2026, DeepSeek grew 12.6x year-over-year; Google grew 3x. US closed-source providers' share of routed requests slid from 54% in May to 44% in late August.

DeepSeek V4 Flash alone burned 45.1T tokens on OpenRouter in August, followed by Tencent's Hy3 at 34.1T and Xiaomi's MiMo-V2.5 at 29.4T. Outside the top ten, Moonshot's Kimi K3 reports 88.3 on Terminal-Bench 2.1, narrowly beating GPT-5.6 Sol's 88.8 on a vendor-run setup.

Open weights, not open source

A quieter table in the report should give pause. Mozilla asked a third party to score 16 marketed open-weight models against the ten OSI criteria; none of them ship a fully disclosed data recipe. Kimi K3 ships under a custom Kimi K3 License with MaaS / UI-attribution restrictions; Qwen3.8-Max under a custom Qwen3.8-Max License with Attribution + MaaS clauses; MiniMax-M3 under a community license with commercial terms. Open weights and open source are still not synonyms, and the training-data side of any major release remains undisclosed.

What closed-source still keeps

The report draws the closed-source edge precisely. Fable 5 leads Kimi K3 by 92 Elo on GDPval-AA v2, the largest Elo gap on any shared benchmark. On 1M-token multi-needle retrieval, Gemini 3.1 Pro reaches 89% versus DeepSeek V4-Pro's 41% — long-context reliability is still a closed-side advantage. SOC 2 / HIPAA / ZDR ship default with the closed-side bundles, and when something goes wrong, having a counterparty to hold accountable is itself a moat.

The headline is not open wins. It is open catches up to 4.4 months on capability and price, while closed keeps the long-context, compliance and liability boundaries. Over the next year, most enterprises will route routine work to open weights and keep the high-stakes workloads on closed APIs. That is exactly the bifurcation Mozilla is betting on.