On September 15 Mozilla published v1.1 of its State of Open Source AI report (first edition July 14, data current to September 1). The headline number: the capability gap between US closed frontier models and the best Chinese open-weight models has narrowed to roughly 4.4 months — Mozilla's own fit on METR time-horizon data, cross-checked against Epoch AI's independent estimate of four months.

From a generation gap to a quarterly one

The report cites the Artificial Analysis Intelligence Index v4.1.1 (September 1 data): in the top ten, the first four seats are closed (Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, Grok 4.6), and the next four are all open weights — Kimi K3 and GLM-5.3 tied at 60.0, three points behind leader Claude Opus 5 at 63.0 and 2.1 points behind Claude Fable 5 at 62.1.

The price comparison is the punchline: Kimi K3 lists at 3/15 USD per million input/output tokens against Fable 5's 10/50 USD — roughly 30% of the cost for a two-point deficit.

Where closed models still earn their premium

The report is candid about what closed frontier still buys. CTO Raffi Krikorian names three areas: expert professional work, high-intensity retrieval, and long context. The hard numbers back him up: Fable 5 leads K3 by 92 Elo on GDPval-AA v2, the largest gap among their shared benchmarks; on 1M-token multi-needle retrieval, Gemini 3.1 Pro scores 89% against DeepSeek V4-Pro's 41%; and on METR's time-horizon data, both open and closed models handle tasks under eight hours, while past twelve hours neither does.

The open-default, closed-on-demand pattern already has a template: DoorDash, cited in the report, runs Kimi for routine workloads and reserves Fable for harder tasks.

Cold water: capability parity, not wallet parity

The sharpest contrast in the report is revenue: in 2025, closed providers captured 96% of model-layer revenue while open models took 4% — even as closed models cost roughly 6x more per call at around 90% parity. Production rates lag too: only 53% of teams running open models get them into production, versus 63% for closed-model teams.

But the balance is moving. August 2026 was the first month an open model led OpenRouter by request count; eight of the top ten models by token volume were open weights, seven of them Chinese-built; US closed providers' request share fell from 54% in May to 44% by late August. One more telling signal: on August 16, DeepSeek delivered the first list-price increase by an open model, raising output prices 2.3-4.6x — when open models gain pricing power, the cheap label deserves a second look.

So what

The real audience for this report is not model vendors but every engineering leader writing an AI budget this quarter: how many of your workloads actually fall inside that four-month head start, and what premium are you paying to cover them? For most organizations the answer is blunt — default to open weights, and reserve closed frontier for the narrow band of tasks that genuinely needs it.

References: Mozilla, State of Open Source AI v1.1; Tom's Hardware coverage