[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mozilla-china-open-weight-ai-gap-4-4-months":3,"topics-all":38,"news-related-c4375463-f274-464e-998e-6f2a9f9cfeee":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"c4375463-f274-464e-998e-6f2a9f9cfeee","Mozilla:中美开放权重AI差距缩至4.4个月","Mozilla 发布第二版《开源 AI 现状》报告,称中国头部开放权重模型与美国前沿闭源模型的能力差距已缩至 4.4 个月。Kimi K3 在 AA 智能指数上仅落后 Fable 5 三分,单任务成本约为后者 30%。八月 token 用量 Top10 中八款为开源权重,其中七款为中国研发。","Mozilla 在 9 月 15 日发布的《State of Open Source AI v1.1》报告把宏观判断落到了具体模型、具体任务时长和单任务成本上。最核心的结论是:中国头部开放权重模型与美国前沿闭源模型的能力差距已经缩至约 4.4 个月。\n\n## 闭源与开源:三分之差,七倍成本\n\n报告引用 Artificial Analysis 智能指数(9 月 1 日版)指出,Anthropic 的 Claude Fable 5 以 62 分位居榜首,Moonshot 的 Kimi K3 以 60 分紧随其后,中间只差三分。但价格层面的差距要大得多:Fable 5 的列表价为 10\u002F50 美元每百万 token,Kimi K3 仅为 3\u002F15 美元——按输入\u002F输出同价计算,Kimi K3 跑完相同任务的有效支出大约只有 Fable 5 的 30%。GLM-5.3 的成本曲线还要更激进:0.075 美元\u002F百万 token 的输入价,基本上把同一档能力的边际成本压到了行业最低。\n\n这就是 Mozilla CTO Raffi Krikorian 给出的「闭源溢价」定性——价格差体现在需要 8 到 12 小时才能完成的专业型任务上,而开源模型则能以极低的价格处理日常工作。\n\n## 任务时长:12 小时仍是分水岭\n\n报告引用 METR Time Horizon 1.1 的方法(任务 50% 可靠成功),把闭源与开源放在了同一个时间轴上:目前最佳闭源模型能可靠完成约 12 小时的任务,最佳开源模型大约能完成 7 小时的任务,任务长度比约为 1.7 倍。Krikorian 的判断是:4 个月后,开源模型就能追上 12 小时,闭源模型则推进到约 20 小时——8 到 12 小时这段中间地带,是当前闭源依然能赚到溢价、开源尚难替代的区间。\n\n这个判断在企业一端已经开始反映在路由选择上。DoorDash 已经把日常任务切到 Kimi,把 Fable 留在更复杂的工作负载里;OpenRouter 8 月按 token 用量统计的前十名里有 8 款是开源权重,其中 7 款为中国团队研发。\n\n## 八月:开源第一次接管请求榜首\n\n报告里最有冲击力的数据点出现在 OpenRouter 这一节。8 月 3 日,DeepSeek 第一次取代 Google 拿下单周请求量第一——Google 已经连续 51 周占据这个位置。从 2025 年 9 月的 800 亿周请求到 2026 年 8 月的 1.01 万亿,DeepSeek 一年的增长倍数是 12.6 倍,Google 的同期增长只有 3 倍。美国闭源头部厂商的请求份额从 5 月的 54% 下滑到 8 月末的 44%。\n\nDeepSeek V4 Flash 单月在 OpenRouter 上消耗 45.1T tokens,其次是腾讯的 Hy3(34.1T)和小米的 MiMo-V2.5(29.4T)。在前十之外,Kimi K3 在自己公布的 Terminal-Bench 2.1 上跑出了 88.3 分,超过了 GPT-5.6 Sol 的 88.8,仅以微弱差距排在开源模型头名。\n\n## OSS vs OSI:十六个「开放」没有真正开源\n\n报告里还有一张容易被忽略的表格:Mozilla 邀请第三方按 OSI 的十条定义重新审视了 16 个所谓「开放权重」模型,结论是没有一个完全满足。Kimi K3 的许可证是定制的「Kimi K3 License」,附 MaaS\u002FUI attribution 限制;Qwen3.8-Max 是阿里定制的「Qwen3.8-Max License」,带 Attribution + MaaS 限制;MiniMax-M3 也只发了社区版许可。换句话说,「开放权重」≠「开放源代码」,真正的训练数据配方,在任何一个主流权重面前都仍然是不公开的。\n\n## 美国闭源的护城河还剩什么\n\n报告把闭源仍然领先的边界画得很具体:Fable 5 在 GDPval-AA v2 上领先 Kimi K3 92 Elo,是所有共享基准里 Elo 差距最大的一个;1M token 多针检索上,Gemini 3.1 Pro 的 89% 对比 DeepSeek V4-Pro 的 41%,长上下文的可靠性差距还很明显。SOC 2 \u002F HIPAA \u002F ZDR 这些企业合规是闭源「打包默认」的卖点;而当一个事故真的发生时,「有责任方可以追责」是另一条护城河。\n\n所以,Mozilla 给出的最终判断不是「开源赢了」,而是「开源在能力与价格上已经追到差 4.4 个月,但在长上下文、企业合规与责任边界上,闭源仍然是工作负载分流的默认选项」。这意味着未来 12 个月里,绝大多数企业的 AI 路线不会变成「all open」,而是「日常上开源,关键任务上闭源」——而这恰好是 Mozilla 想要看到的分流。","https:\u002F\u002Fstateofopensource.ai\u002Fstate-of-open-source-ai-v1-1.pdf","8a4f9b4d-8163-4866-93d0-3611ef77a0bb",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"5ea844e8-2ab7-4751-b094-f6eaa3e848d7","en","Mozilla: US-China open-weight AI gap narrows to 4.4 months","Mozilla published the second edition of the State of Open Source AI report on September 15, narrowing the capability gap between Chinese open-weight models and US frontier closed-source models to roughly 4.4 months. Kimi K3 trails Claude Fable 5 by only three points on the AA Intelligence Index while costing about 30% of the latter. In August, eight of the top ten models by token volume on OpenRouter were open-weight, with seven built by Chinese teams.","Mozilla published version 1.1 of the State of Open Source AI report on September 15, and the headline number is sharp: the capability gap between Chinese open-weight models and US frontier closed-source models has compressed to about 4.4 months.\n\n## Three points on the leaderboard, seven times the cost\n\nDrawing on the September 1 cut of the Artificial Analysis Intelligence Index, Anthropic's Claude Fable 5 sits on top at 62 points, with Moonshot's Kimi K3 at 60. The score gap is just three points. The price gap is far wider. Fable 5 lists at $10 input \u002F $50 output per million tokens; Kimi K3 lists at $3 \u002F $15. Run the same workload through both, and Kimi K3's effective spend lands near 30% of Fable 5's. Z.ai's GLM-5.3 pushes the cost curve further down: $0.075 input per million tokens, the cheapest seat at this capability tier.\n\nMozilla CTO Raffi Krikorian frames that gap as a premium that lives almost entirely in the 8-to-12-hour professional task band. Below that, open weights do the job at a fraction of the cost.\n\n## Twelve hours is still the closed-source moat\n\nCiting METR Time Horizon 1.1 (50% reliable-success criterion), the report lines both sides up on the same time axis: the best closed-source model now reliably completes a ~12-hour task; the best open-weight model reaches about 7 hours, a task-length ratio of roughly 1.7x. Krikorian's call is that four months from now, open weights catch 12 hours while closed-source pushes to ~20 hours, meaning the 8-to-12-hour band is exactly where closed-source still earns its premium today.\n\nThat call is already showing up in routing decisions. DoorDash has moved day-to-day workloads to Kimi while reserving Fable for more complex jobs. On OpenRouter's August token-volume ranking, eight of the top ten models are open-weight, and seven of the eight are built by Chinese teams.\n\n## August was the first month an open-weight model led OpenRouter\n\nThe most striking data point in the report sits in the OpenRouter section. On August 3, DeepSeek took over the weekly-request crown from Google, which had held it for 51 straight weeks. From 80 billion weekly requests in September 2025 to 1.01 trillion in August 2026, DeepSeek grew 12.6x year-over-year; Google grew 3x. US closed-source providers' share of routed requests slid from 54% in May to 44% in late August.\n\nDeepSeek V4 Flash alone burned 45.1T tokens on OpenRouter in August, followed by Tencent's Hy3 at 34.1T and Xiaomi's MiMo-V2.5 at 29.4T. Outside the top ten, Moonshot's Kimi K3 reports 88.3 on Terminal-Bench 2.1, narrowly beating GPT-5.6 Sol's 88.8 on a vendor-run setup.\n\n## Open weights, not open source\n\nA quieter table in the report should give pause. Mozilla asked a third party to score 16 marketed open-weight models against the ten OSI criteria; none of them ship a fully disclosed data recipe. Kimi K3 ships under a custom Kimi K3 License with MaaS \u002F UI-attribution restrictions; Qwen3.8-Max under a custom Qwen3.8-Max License with Attribution + MaaS clauses; MiniMax-M3 under a community license with commercial terms. Open weights and open source are still not synonyms, and the training-data side of any major release remains undisclosed.\n\n## What closed-source still keeps\n\nThe report draws the closed-source edge precisely. Fable 5 leads Kimi K3 by 92 Elo on GDPval-AA v2, the largest Elo gap on any shared benchmark. On 1M-token multi-needle retrieval, Gemini 3.1 Pro reaches 89% versus DeepSeek V4-Pro's 41% — long-context reliability is still a closed-side advantage. SOC 2 \u002F HIPAA \u002F ZDR ship default with the closed-side bundles, and when something goes wrong, having a counterparty to hold accountable is itself a moat.\n\nThe headline is not open wins. It is open catches up to 4.4 months on capability and price, while closed keeps the long-context, compliance and liability boundaries. Over the next year, most enterprises will route routine work to open weights and keep the high-stakes workloads on closed APIs. That is exactly the bifurcation Mozilla is betting on.","mozilla-china-open-weight-ai-gap-4-4-months","2026-09-26T00:00:00Z","2026-09-26T01:03:45.335877Z","2026-09-26T01:03:45.335892Z",true,"agent",390,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"9389d1ed-dd2d-41cb-bbc5-9a543e2b2f71","开源报告里的「参数天花板」分水岭:中国实验室把上限拉到2.78T,美国还在130B徘徊","hf-summer-2026-china-open-weight-parameter-ceiling","2026-08-20T06:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"4bb31ede-b9c4-4762-86ae-9d3b008557ca","Hugging Face Summer 2026 报告:Qwen 拿下 15 万衍生模型, GGUF 仓库一年涨 464%","hugging-face-state-of-open-models-summer-2026","2026-08-18T02:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"5314fe6d-ba17-42bc-9f52-197b8cb9cf91","黄仁勋力挺中国开源大模型:中美技术差距共识正在被开源生态改写","jensen-huang-china-open-source","2026-07-24T03:35:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"c8b1fd9a-524e-4037-adea-d850696291c2","微软测试 DeepSeek V4 接入 Copilot：开源 LLM 首次威胁到头部办公软件的核心","microsoft-copilot-deepseek-v4-open-source","2026-06-22T16:30:00+00:00"]