In early September, the core team behind Wagtail, the open-source CMS, set itself a challenge: for one whole month, all engineering work would run on a single efficient open-weight model, GLM-5.3-Flash. On October 2, team member Thibaud Colas published the retrospective, subtitled with a piece of self-deprecating humor: "task failed successfully". Two billion tokens later, the single-model goal was only half met.
Where the 2B tokens went
According to AgentsView, the usage-tracking tool the team recommends, the first half of the month ran 100% on GLM-5.3-Flash. That stretch cost 68 dollars, roughly 4kWh of energy and 365 grams of CO2 - strikingly cheap for a month of AI-assisted engineering. The second half went sideways: another 1B tokens flowed to other models, and total energy use landed around 35kWh, roughly 3.5 times what they had budgeted.
Three cracks
First, the cost of vibe coding. The experimental Wagtail MCP server prototype was built with the wrong model, and burned 450M tokens, 150 dollars and 5kWh almost overnight. The team's own estimate: with better model selection, the same result could most likely have cost about a fifth as much.
Second, inference infrastructure wobbled. The third-party inference providers they rely on ran into capacity problems, GLM-5.3-Flash performance degraded, and the team switched to DeepSeek V4.1 Flash and Qwen 3.8 Flash. Their explanation is blunt: the model sits so high on the Pareto frontier of models relevant to their work that everyone piles onto it, and these providers do not have the GPU capacity of the big labs who hoard them all.
Third, experimentation has its own budget. Building a benchmark of models on Wagtail's own tasks requires coverage across many models by definition - those tokens can never count toward a single-model goal.
The 14-model benchmark, as a bonus
Perhaps the most valuable byproduct of the retrospective is a preview of a benchmark table: 14 models scored on Wagtail tasks by accuracy, energy use and cost. At the top sits DeepSeek V4.1 Flash - 95% accuracy, 14.9Wh and 0.09 dollars per task.
So what
Three observations. First, the failure deserves air quotes: the team's own conclusion is that day-to-day development work is perfectly viable on one or two cheap flash-tier models - what failed was the more radical single-model constraint. Second, the bottleneck is shifting from "is the model good enough" to "is there enough inference capacity": when one cheap, capable open model becomes everyone's default, the scaling speed of third-party inference services becomes the new weak link. Third, a detail easy to miss: the protagonist GLM-5.3-Flash plus both fallback models all come from Chinese teams - Chinese open-weight models running the daily workflow of a Western open-source engineering team is no longer a talking point, it is a line item in the ledger.
For October, the team is switching how it measures: not tokens, but cost and energy, with a goal of letting flash-tier models absorb the majority of inference work. That playbook is worth stealing for any team drawing up an AI budget.
References: Wagtail's retrospective and the September challenge announcement.