[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cognition-swe-2-kimi-k3-pareto":3,"topics-all":35,"news-related-01a6593b-449e-493e-ac45-33c23c9211ba":54},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"01a6593b-449e-493e-ac45-33c23c9211ba","SWE-2 距 Fable 5.1 一分:2.8T 开源底座后训练,成本砍 64%","Cognition 发布编码模型 SWE-2:以 2.8T 开放的 Kimi K3 为底座做 RL 后训练,FrontierCode 1.1 Main 得 50.0%、距 Fable 5.1 仅 0.9 分且便宜 64%;单次 RL 同时训完三档推理档位,但 Terminal-Bench 4 落后前沿约 28 分。","50.0% 对 50.9%,差 0.9 个百分点,价格差 64%。这是 Cognition 9 月 10 日发布的编码模型 SWE-2 交出的成绩单:在自家的 FrontierCode 1.1 Main 上,它把自己和 Anthropic 的 Fable 5.1 摆到了几乎同一条线上,而成本只有对方的零头。\n\n## 开源底座的\"再训练\"红利\n\nSWE-2 的底座不是自研基座,而是月之暗面 2.8T 参数的开放权重模型 Kimi K3——一个已经为 agentic coding 做过大量 RL 的模型。Cognition 在官方博文中称,他们的 RL 配方在这个\"熟模型\"上仍然找到了可观余量:多个基准上再涨 5-6 分,整条成本-性能前沿被平移。按官方说法,这也是该团队首次把 RL 后训练扩到万亿参数量级。\n\n方法论上真正的看点,是单次 RL 训练同时覆盖全部推理档位。奖励函数写作 R = S − λe·C:S 是任务成败,C 是 rollout 成本(美元推理费与时间的混合),λe 则按基座 Pareto 前沿在每个档位的切线斜率去设。直观效果:模型不会被诱导\"降档偷懒\"——高努力档位假装成中档可以省成本,但 reward 不会因此增加,只有真实推高前沿才会。配套工程同样扎实:长度加权 reward baseline 稳住训练 KL;DSpark 投机解码加 SpecForge 训出的新 draft 模型把接受长度拉长 15%,并在线持续追踪策略变化;NVFP4\u002FFP8 内核加量化感知训练压住显存。行为变化很直观:SWE-2 medium 做出首次真实编辑的中位步数从 SWE-1.7 的 48 步降到 18 步,平均轮次少 58%,均费便宜 81%——过度探索的毛病被\"会判断哪段代码值得读\"替代了。\n\n## 成绩单与那盆冷水\n\n官方表格四行基准:FrontierCode 1.1 Main 50.0%(Kimi K3 44.2、Grok 4.6 48.0、GPT-5.6 Sol 47.5、Fable 5.1 50.9、GPT-6 Astra 53.3、SWE-1.7 42.0);DeepSWE 1.1 拿 73.0%,仅次于 GPT-6 Astra 的 74.1%;Terminal-Bench 2.1 得 92.8%,是唯一一行压过表内全部对手的成绩。冷水在第四行:Terminal-Bench 4 只有 27.3%,而 Fable 5.1 是 55.8%、GPT-6 Astra 57.9%——第三方分析直言,\"与前沿打平\"在前三行成立,第四行不成立。长程终端任务上的这个差距,是\"便宜 64%\"叙事里最该被记住的脚注。\n\n## 所以呢\n\nSWE-2 已上线 Devin Desktop 与 CLI,并陆续铺到 Devin Web 和 Fusion。对行业,这件事的分量不在某一行跑分,而在它再次验证了一条路径:开放权重的价值不只是免费用,而是可以被第三方用一套公开方法论后训练成逼近前沿的商用模型。当 2.8T 的开源底座,遇上会训练整条 Pareto 前沿的团队,闭源厂商的溢价空间还剩多少?\n\n参考:[Cognition 官方博文](https:\u002F\u002Fcognition.com\u002Fblog\u002Fswe-2);cellcog.ai 基准拆解;AI\u002FTLDR 摘要。","https:\u002F\u002Fcognition.com\u002Fblog\u002Fswe-2","0413dae1-4326-4b2b-b457-ee8c2ee5485f",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"efce2154-58d0-4cf5-9461-6e0430f2e631","en","SWE-2 lands within a point of Fable 5.1 at 64% lower cost","SWE-2: RL post-train of open 2.8T Kimi K3 hits 50.0% FrontierCode, 0.9 pt behind Fable 5.1 at 64% less cost; one RL run trains all effort levels.","50.0% versus 50.9% — a 0.9-point gap, at 64% lower cost. That is the headline from Cognition's September 10 release of SWE-2: on its own FrontierCode 1.1 Main benchmark, the model sits on nearly the same line as Anthropic's Fable 5.1 while costing a fraction of the price.\n\n## The re-training dividend of an open base\n\nSWE-2 is not built on an in-house foundation model. Its base is Kimi K3, Moonshot AI's 2.8-trillion-parameter open-weight model — one that had already undergone extensive RL for agentic coding. According to Cognition's official post, its RL recipe still found substantial headroom on this already-trained model: 5–6 additional points on many benchmarks, shifting K3's entire cost–performance frontier. By the company's own account, this is also the first time the team has scaled RL post-training into the multi-trillion-parameter regime.\n\nThe methodological highlight is that a single RL run now covers every reasoning-effort level. The reward reads R = S − λe·C, where S is task success, C is rollout cost (a mix of inference cost in USD and time), and λe is tuned to the slope of the base model's Pareto frontier at each effort level. The practical consequence: the model cannot profit by \"downshifting\" — at high effort, pretending to be the medium tier saves cost but does not increase reward; only genuinely pushing the frontier up does. The surrounding engineering is equally dense: a length-weighted reward baseline keeps training KL stable; DSpark speculative decoding plus a SpecForge-trained draft model lengthens accepted sequences by 15% while tracking the policy online; NVFP4\u002FFP8 kernels with quantization-aware training rein in memory. The behavioral shift is visible: SWE-2 medium makes its first real edit after a median of 18 steps versus 48 for SWE-1.7, takes 58% fewer turns, and costs 81% less on average — over-exploration is replaced by knowing which part of the codebase actually matters.\n\n## The scorecard, and the cold water\n\nThe official table lists four benchmarks. FrontierCode 1.1 Main: 50.0% (Kimi K3 44.2, Grok 4.6 48.0, GPT-5.6 Sol 47.5, Fable 5.1 50.9, GPT-6 Astra 53.3, SWE-1.7 42.0). DeepSWE 1.1: 73.0%, second only to GPT-6 Astra's 74.1%. Terminal-Bench 2.1: 92.8% — the only row where SWE-2 beats every model listed. The cold water sits in the fourth row: Terminal-Bench 4 comes in at 27.3%, against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. Independent analysis states plainly that \"on par with the frontier\" holds on three benchmarks and not the fourth; that long-horizon terminal gap is the footnote most worth remembering inside the \"64% cheaper\" story.\n\n## So what\n\nSWE-2 is live in Devin Desktop and CLI, and rolling out to Devin Web and Fusion. For the industry, the significance is less any single score than the path it re-validates: the value of open weights is not just free access, but that a third party can post-train them into a commercial model that closes on the frontier with a published methodology. When a 2.8T open base meets a team that knows how to train the whole Pareto frontier, how much pricing power do closed labs have left?\n\nReferences: [Cognition's official post](https:\u002F\u002Fcognition.com\u002Fblog\u002Fswe-2); cellcog.ai benchmark breakdown; AI\u002FTLDR digest.","cognition-swe-2-kimi-k3-pareto","2026-09-11T21:08:02Z","2026-09-11T21:09:11.720724Z","2026-09-11T21:09:11.720734Z",true,"agent",56,[36,45],{"slug":37,"tag_slug":37,"title_zh":38,"title_en":39,"intro_zh":40,"intro_en":41,"id":42,"is_active":32,"created_at":43,"modified_at":44},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":32,"created_at":52,"modified_at":53},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":55},[56,61,66,71,76,81],{"id":57,"title":58,"news_slug":59,"published_at":60},"8ebbcd9c-31ee-4baa-b395-b104bd87c8e1","Kimi K2.8 Preview 把 K3 的百万上下文下放给免费档：月之暗面的「过日子」模型登场","kimi-k2-8-preview-coding","2026-09-17T03:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"d400c0db-49df-4cc6-a87e-87b709f59fea","Muse Spark 1.3 发布:卡住会向用户求助的 Agent,工具调用少 20%、token 省 25%","muse-spark-1-3-meta-agent-release","2026-09-06T15:12:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"89a79f9a-bfd2-4ebe-8f03-92fa74a3a34f","Ornith-1.5 开源：模型自己出题、自己搭考场，397B 到 9B 三档齐发","ornith-1-5-self-improvement-open-models","2026-08-20T13:30:00+00:00",{"id":72,"title":73,"news_slug":74,"published_at":75},"f8a3ead9-b1aa-4f6a-9146-19c01f7f1375","微软把编码模型价格砍到四分之一：138B\u002F5B 稀疏 MoE 加原生视觉全量进入 Copilot","mai-code-1-1-flash-copilot-moe-vision","2026-08-12T06:00:00+00:00",{"id":77,"title":78,"news_slug":79,"published_at":80},"e846f1c6-4644-4f84-a664-83ec82734210","Meta Muse Spark 1.2 与 Muse Code 把「1.2 → 编程」的推理效率推回前沿","meta-muse-spark-12-coding-agent-54-index","2026-08-05T08:00:00+00:00",{"id":82,"title":83,"news_slug":84,"published_at":85},"dea38861-2618-4468-9bad-a18eea96a818","Base44 Base 1：年入 1.5 亿美元的 vibe-coding 平台，终于把自己的 LLM 训出来","base44-base-1-vibe-coding-llm-launch","2026-07-29T06:00:00+00:00"]