[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wagtail-glm-5-3-flash-month":3,"topics-all":38,"news-related-474e602e-a509-4973-8d44-eafc532138a4":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"474e602e-a509-4973-8d44-eafc532138a4","GLM-5.3-Flash 一个月:Wagtail 半程失守","开源 CMS Wagtail 团队九月实验:全月工程任务只押 GLM-5.3-Flash。2B token 账本:前半月纯跑 68 美元约 4kWh 电,后半月推理容量退化,一半 token 改投 DeepSeek 与 Qwen。团队结论:日常开发押一两个 flash 档便宜模型完全可行。","9 月初,开源 CMS Wagtail 的核心团队给自己出了道题:整个月的工程任务,只准用 GLM-5.3-Flash 这一个开放权重模型。10 月 2 日,团队成员 Thibaud Colas 交出复盘,副标题是句自嘲——「task failed successfully」,任务失败了,但失败得信息量十足:一个月 2B token 花出去,单模型目标只完成了一半。\n\n## 账本:2B token 花在哪了\n\n团队的 AI 用量统计工具 AgentsView 给出的账目是:前半个月,他们 100% 跑在 GLM-5.3-Flash 上,这部分成本 68 美元、约 4kWh 电、365 克碳排放——对一整个月的 AI 辅助工程来说,这个数字便宜得有些反常。转折在后半月:另外 1B token 流向了其他模型,全月能耗约 35kWh,是他们原本预期的三倍半。\n\n## 三条裂缝\n\n**一是 vibe coding 的代价。** 实验性的 Wagtail MCP 服务器原型选错了模型,几乎一夜之间烧掉 450M token、150 美元、5kWh 电。团队自己估算,选对模型的话,同样的结果成本大概率只要五分之一。\n\n**二是推理基础设施掉链子。** 他们依赖的第三方推理供应商容量不足,GLM-5.3-Flash 的服务质量出现退化,后半程被迫切换到 DeepSeek V4.1 Flash 和 Qwen 3.8 Flash。团队的解释相当直白:这个模型在他们画出的性价比 Pareto 前沿上位置太靠前、太受欢迎,而供应商的 GPU 家底远不如囤卡的大厂,扛不住这么多人同时挤上来。\n\n**三是实验预算天然超支。** 团队要给 Wagtail 自家任务做模型基准,基准研究本身就必须覆盖多家模型,这部分 token 无论如何记不进单模型目标的账。\n\n## 顺手晒出的 14 模型基准\n\n复盘里最值钱的附产品,可能是一张 14 个模型在 Wagtail 任务上的基准表预览:准确率、能耗、成本三列并排。榜首是 DeepSeek V4.1 Flash——95% 准确率、单任务 14.9Wh、0.09 美元。\n\n## 所以呢\n\n三点观察。第一,「失败」要打引号:团队自己的结论是,日常开发工作完全可以在一两个 flash 档便宜模型上运转——被验证失败的,是「单模型」这个更激进的目标。第二,瓶颈正在从「模型行不行」挪到「推理容量够不够」:当一个便宜好用的开放模型被所有人同时盯上,第三方推理服务的扩容速度就成了新的短板。第三,一个容易被忽略的细节:主角 GLM-5.3-Flash 加上两个备选,清一色出自中国团队——国产开放模型跑进西方开源工程团队的日常工作流,已经不是口号,而是账本上的既成事实。\n\n10 月,这个团队换了计量方式:不再数 token,改按成本和能耗算账,目标是让 flash 档模型吃下多数推理任务。这套思路,值得所有正在做 AI 预算的团队抄作业。\n\n参考:[Wagtail 团队复盘原文](https:\u002F\u002Fwagtail.org\u002Fblog\u002Fone-month-on-glm-53-flash\u002F);[九月挑战公告](https:\u002F\u002Fwagtail.org\u002Fblog\u002Fopen-models-only-adoption-challenge\u002F)。","https:\u002F\u002Fwagtail.org\u002Fblog\u002Fone-month-on-glm-53-flash\u002F","cc95e520-b8a0-4c83-9013-7764f7e657bb",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"43ad222b-3902-4c42-ad6b-61e28dc2cd26","en","One month on GLM-5.3-Flash: Wagtail's 2B-token verdict","Wagtail spent September on one model: $68\u002F4kWh covered half a month on GLM-5.3-Flash; capacity issues pushed 1B of 2B tokens to DeepSeek and Qwen.","In early September, the core team behind Wagtail, the open-source CMS, set itself a challenge: for one whole month, all engineering work would run on a single efficient open-weight model, GLM-5.3-Flash. On October 2, team member Thibaud Colas published the retrospective, subtitled with a piece of self-deprecating humor: \"task failed successfully\". Two billion tokens later, the single-model goal was only half met.\n\n## Where the 2B tokens went\n\nAccording to AgentsView, the usage-tracking tool the team recommends, the first half of the month ran 100% on GLM-5.3-Flash. That stretch cost 68 dollars, roughly 4kWh of energy and 365 grams of CO2 - strikingly cheap for a month of AI-assisted engineering. The second half went sideways: another 1B tokens flowed to other models, and total energy use landed around 35kWh, roughly 3.5 times what they had budgeted.\n\n## Three cracks\n\n**First, the cost of vibe coding.** The experimental Wagtail MCP server prototype was built with the wrong model, and burned 450M tokens, 150 dollars and 5kWh almost overnight. The team's own estimate: with better model selection, the same result could most likely have cost about a fifth as much.\n\n**Second, inference infrastructure wobbled.** The third-party inference providers they rely on ran into capacity problems, GLM-5.3-Flash performance degraded, and the team switched to DeepSeek V4.1 Flash and Qwen 3.8 Flash. Their explanation is blunt: the model sits so high on the Pareto frontier of models relevant to their work that everyone piles onto it, and these providers do not have the GPU capacity of the big labs who hoard them all.\n\n**Third, experimentation has its own budget.** Building a benchmark of models on Wagtail's own tasks requires coverage across many models by definition - those tokens can never count toward a single-model goal.\n\n## The 14-model benchmark, as a bonus\n\nPerhaps the most valuable byproduct of the retrospective is a preview of a benchmark table: 14 models scored on Wagtail tasks by accuracy, energy use and cost. At the top sits DeepSeek V4.1 Flash - 95% accuracy, 14.9Wh and 0.09 dollars per task.\n\n## So what\n\nThree observations. First, the failure deserves air quotes: the team's own conclusion is that day-to-day development work is perfectly viable on one or two cheap flash-tier models - what failed was the more radical single-model constraint. Second, the bottleneck is shifting from \"is the model good enough\" to \"is there enough inference capacity\": when one cheap, capable open model becomes everyone's default, the scaling speed of third-party inference services becomes the new weak link. Third, a detail easy to miss: the protagonist GLM-5.3-Flash plus both fallback models all come from Chinese teams - Chinese open-weight models running the daily workflow of a Western open-source engineering team is no longer a talking point, it is a line item in the ledger.\n\nFor October, the team is switching how it measures: not tokens, but cost and energy, with a goal of letting flash-tier models absorb the majority of inference work. That playbook is worth stealing for any team drawing up an AI budget.\n\nReferences: [Wagtail's retrospective](https:\u002F\u002Fwagtail.org\u002Fblog\u002Fone-month-on-glm-53-flash\u002F) and [the September challenge announcement](https:\u002F\u002Fwagtail.org\u002Fblog\u002Fopen-models-only-adoption-challenge\u002F).","wagtail-glm-5-3-flash-month","2026-10-03T17:08:59Z","2026-10-03T17:10:01.385906Z","2026-10-03T17:10:01.385915Z",true,"agent",820,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"3d8b9b1a-e038-466f-9b6b-304f911e35a7","Kimi K3 开源三件套 MoonEP\u002FFlashKDA\u002FAgentEnv:Moonshot 把 2.8T MoE 训练栈完整交底","kimi-k3-moonep-flashkda-agentenv","2026-07-28T04:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"4ccc491b-dbee-4c84-beb6-7cf519f76320","LoRA 基座换 GGUF:40G 显存训 125B","lora-over-gguf-low-vram-training","2026-10-10T21:08:22+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"e304fe91-16d5-4c28-81c7-0c4c0dd8c669","小米公开MiMo-V2.6训练账本:RL烧了260万美元","xiaomi-mimo-v2-6-rl-scaling-ledger","2026-10-10T13:11:49+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"530e5aa5-2026-4c30-b4e6-421caca907b2","Transformer 提前罢工:13 个基座模型跟不住引用链,一个 rank-8 LoRA 修好","tiny-lora-frozen-transformer-chain-relay","2026-10-03T21:05:16+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"9a77e0d7-8164-494e-984d-54cb57a7a0dd","Magnitude 开源:本地 Agent 专用推理引擎","magnitude-self-tuning-local-agent-inference-engine","2026-10-01T17:08:35+00:00"]