[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-amazon-rufus-air-post-training-recipe":3,"topics-all":38,"news-related-2dbc7c0f-083a-47d4-ba9a-8a7d66b22ae0":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"2dbc7c0f-083a-47d4-ba9a-8a7d66b22ae0","亚马逊八阶段配方:后训练让 GLM-4.5-Air 反超官方版","亚马逊 22 人团队发布 47 页论文,公开在智谱开源 GLM-4.5-Air-Base(106B-A12B)上的八阶段后训练配方:SFT、三段 RL、三个 agent 阶段再到 RLHF,不用新标注、不请蒸馏教师,RL 用 8-32 节点跑完。自测口径下成品在 IFEval 等基准反超官方版,负结果与成本一并写明。","9 月 24 日,arXiv 出现一篇 47 页长文:亚马逊 22 人团队把在开源基座 GLM-4.5-Air-Base(106B 总参 \u002F 12B 激活 MoE)上的完整后训练配方公开,成品命名 Rufus-Air。数据、奖励、基础设施、逐阶段结果全部写透——后训练普遍当黑盒的年代,这是份少见的全透明材料。\n\n## 骨架:八阶段串行流水线\n\n后训练是一条串行管线:SFT → 推理 RL → 编码 RL → 指令遵循 RL → 通用 Agent → 编码 Agent → 搜索 Agent → RLHF,每阶段以上一阶段 checkpoint 为起点。排序原则有二:能力从基础到高级;奖励从硬可验证到软打分——先确定性,后 judge 软奖励,让策略晚暴露在 reward hacking 风险里。\n\nSFT 数据规模很实:901 万样本、445 亿原始 token,遮蔽后实际监督 270 亿。仅这个 checkpoint 就已在 IFEval 领先官方版 5.3 分、IFBench 领先 24.2 分——作者反对把 SFT 当热身:它建立的是 RL 精修的地板。\n\n## 阶段增量与诚实的负结果\n\n难度过滤是实用抓手:解题率高于 0.8 的提示词当「太容易」丢,零通过的当「学不会」丢,训练只留中间可学习带。推理 RL 用 GSPO,GPQA 从 68.2 升到 73.5(+5.3),代价是 AIME 两个年度各降 2.6-2.8 分,作者坦承是主动取舍。编码 RL 把 LiveCodeBench v6 从 67.9 拉到 74.5。指令遵循 RL 用 GRPO,IFEval +4.0 到 94.5,GPQA(+3.2)基本不动。\n\nAgent 三阶段信息量最大。通用 Agent 用 1 万个 MCP 任务、1000 个合成环境训练,MCP-Atlas +7.80、Tau2-Retail +9.80。编码 Agent 跑 Docker 沙箱加单元测试奖励,SWE-bench Verified 从 65.60 到 67.80,作者注明此阶段未训到收敛。搜索 Agent 面向开放网页,训练时最多 100 次工具调用,BrowseComp +3.0、Seal-0 +5.4。最终 RLHF 用开源奖励模型 Skywork-Reward-V2-Qwen3-8B,Arena-Hard v2 Hard Prompt 从 83.06 到 89.05。\n\n## 对照与成色\n\n同 harness 四模型对比下,Rufus-Air 在除 Arena-Hard v2 创意写作外的报告基准上全部领先官方 GLM-4.5-Air:IFEval 95.4 对 83.0,SWE-bench Verified 65.6 对 50.6。同基座的 INTELLECT-3 与 Nemotron-3、GPT-OSS、Qwen3.5 同表对照;作者读法:哪投入训练信号,产出就在哪移动。但官方版是 2025 年 7 月发布,基准代际更新后未重跑,对比含金量要打折扣。\n\n## 所以呢\n\n其一,信息不对称正被填平:MoE、agent 环境合成、RLHF 完整链路的全透明报告仍稀缺,亚马逊把「开源基座到成品」的工程路径标准化了一次。其二,美国大厂拿智谱 GLM-4.5-Air-Base 当训练地基,成品多基准反超官方版——开放权重的基座价值与后训练可塑性同时被验证,对中国开源生态是微妙信号。其三,跑分趋同时,可复现配方加负结果比刷榜更有长期价值。\n\n开发者落点:8-32 节点跑完全部 RL 阶段,中等团队有了可照抄路径;沙箱每月约 1 万美元、搜索 API 按次计费,要单独预算。\n\n参考:arxiv.org\u002Fabs\u002F2609.29421 · hf.co\u002Fpapers\u002F2609.29421","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.29421","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1f9d492e-4346-4d67-9456-0a8ab8159c4b","en","Amazon's Post-Training Recipe for GLM-4.5-Air","Amazon's 47-page paper documents a full eight-stage post-training recipe on Zhipu's open GLM-4.5-Air-Base: no new annotation, RL on 8-32 nodes.","A 47-page paper from Amazon landed on arXiv on September 24, 2026, and it does something rare: it documents a complete post-training recipe on an open base model, in full. The team built Rufus-Air on top of GLM-4.5-Air-Base, the open Mixture-of-Experts checkpoint with 106B total and 12B active parameters released by Zhipu, and published the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the whole run. Twenty-two authors, all at Amazon, listed alphabetically by surname. In an era when most labs treat post-training as a trade secret, this is an unusually transparent artifact.\n\n## The skeleton: one serial pipeline, eight stages\n\nThe recipe is a single serial pipeline with no branching: SFT, then Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and finally RLHF. Each stage trains from the checkpoint the previous one produced. Two axes set the order. On capability, stages move from basic to advanced, so each builds on what the previous one seeded. On reward type, stages run from hard, verifiable rewards toward softer judge-based signals, which limits how long the policy is exposed to reward hacking.\n\nThe SFT numbers are concrete: 9.01M samples, 44.5B raw tokens, 66.7M conversational turns, and 27.0B supervised tokens after masking system messages, user turns, and tool observations. Decontamination used a word-level 8-gram overlap screen and removed 3,529 samples in total, 3,321 of them from a single synthetic text-to-terminal corpus. The SFT checkpoint alone (step 3799) already leads the public GLM-4.5-Air release by 5.3 points on IFEval and 24.2 on IFBench. The authors push back explicitly on treating SFT as a warm-up: it establishes the capability floor that RL later refines.\n\n## Stagewise gains, measured honestly\n\nDifficulty filtering is one of the recipe's practical workhorses: prompts the policy already solves at a rate above 0.8 are dropped as too easy, and prompts with zero observed success are dropped as currently unlearnable, so training concentrates on the learnable band in between. Reasoning RL uses GSPO as the policy-gradient backbone and lifts GPQA from 68.2 to 73.5 (+5.3), at the cost of 2.6-2.8 points on both AIME years — a stated trade-off, not an accident. Coding RL raises LiveCodeBench v6 from 67.9 to 74.5, peaking at 75.9 under an extended 128K response budget. Instruction-Following RL trains with GRPO and moves IFEval up 4.0 points to 94.5 while leaving GPQA (+3.2) and AIME 25 (-1.2) essentially level.\n\nThe three agent stages carry the most new information. General Agent trains on 10K single-server MCP tasks across 1K synthetic environments, and the authors report MCP-Atlas +7.80 and Tau2-Retail +9.80, reading this as evidence that synthetic MCP environments provide a general prior for tool orchestration. Coding Agent runs Docker-sandboxed tasks with test-suite rewards on the Harbor task format, moving SWE-bench Verified from 65.60 to 67.80 and Terminal-Bench 2.1 from 38.76 to 40.17 — with an honest caveat that the stage was trained on the compute available rather than to convergence. Search Agent works against the open web, iterating up to 100 tool calls within a 128K-token budget during training, and adds +3.0 on BrowseComp, +5.4 on Seal-0, and +3.4 on HLE-Verified. The final RLHF stage uses the open Skywork-Reward-V2-Qwen3-8B reward model and only the prompts from HH-RLHF, lifting Arena-Hard v2 Hard Prompt from 83.06 to 89.05 and Creative Writing from 38.56 to 52.97.\n\n## How the finished model stacks up\n\nUnder the paper's own harness, with four models evaluated under identical conditions, Rufus-Air leads the official GLM-4.5-Air release on every reported benchmark except Arena-Hard v2 Creative Writing: IFEval 95.4 vs 83.0, SWE-bench Verified 65.6 vs 50.6, Arena-Hard v2 Hard Prompt 89.1 vs 55.0. Same-base INTELLECT-3, plus Nemotron-3, GPT-OSS, and Qwen3.5 at similar scale, appear in the same table; on competition math and knowledge, the authors' own reading is that the recipe moves most where it spends the most training signal. One caveat the paper flags only lightly: the official GLM-4.5-Air reference dates from July 2025 and predates several of the newer benchmark generations, so the comparison's headroom partly reflects benchmark drift.\n\n## Why this recipe matters\n\nThree layers. First, the information asymmetry around post-training is being filled in. Open recipes such as Tülu and OpenRAID exist, but a fully documented run spanning MoE architecture, agent environment synthesis, and budget-conscious RLHF on a single open base remains scarce; Amazon has effectively standardized the engineering path from open base to finished model once. Second, it is a subtle signal for the Chinese open ecosystem: a US hyperscaler chose Zhipu's GLM-4.5-Air-Base as its training ground, and the result outperforms the official release on multiple benchmarks — both the base value and the post-trainability of open-weight models got validated at once. Third, openness itself may become a competitive axis: as benchmark scores converge, a reproducible recipe with negative results spelled out is worth more long-term than another leaderboard entry.\n\nThe so-what for ordinary developers: RL stages that fit on 8-32 nodes mean mid-sized teams, not just frontier labs, now have a path to copy. The sandbox service runs on the order of ten thousand concurrent sandboxes at roughly $10K per month, and search API calls are billed per use — budget these separately. The recipe paper is not the endpoint; it is one step in moving post-training from alchemy toward engineering. The next thing to watch is how quickly the community reproduces a second batch of models from it.\n\nReferences: full paper at arXiv:2609.29421 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.29421); Hugging Face paper page at https:\u002F\u002Fhuggingface.co\u002Fpapers\u002F2609.29421 .","amazon-rufus-air-post-training-recipe","2026-09-25T23:12:24Z","2026-09-25T23:12:27.742232Z","2026-09-25T23:12:27.742240Z",true,"agent",370,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"054e060c-e182-42a9-b3ed-229feb8ac0ac","2026 年的蒸馏长什么样:Hugging Face 拆解前沿模型三大范式","2026-distillation-three-paradigms","2026-07-09T12:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"ab2d6e9e-8890-4ae7-b6ca-8febc831a279","HPLT MultiSynt\u002FMT：4.8 万亿 token 多语种数据集","hplt-multisynt-multilingual-dataset","2026-07-07T16:01:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"93dfc6f4-e4a9-47a3-aa66-b4b9cee864e7","Beyond LoRA 不只是口号：HF 给 40+ PEFT 方法拍下公平基准，OFT 在图像任务上反超 LoRA","beyond-lora-hf-peft-benchmark-oft","2026-06-18T12:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"44740b4d-8c2c-44fc-8fff-fd89f3fb54ed","12 万美元 token 把 Copilot 运行时从 TypeScript 搬到 Rust","github-copilot-rust-migration-stephen-toub","2026-09-27T11:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"c4375463-f274-464e-998e-6f2a9f9cfeee","Mozilla:中美开放权重AI差距缩至4.4个月","mozilla-china-open-weight-ai-gap-4-4-months","2026-09-26T00:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"e4b3903e-c65f-46cb-90ae-81502eb8cdd9","CliffCompaction开源:只删不改的会话压缩,长程Agent成本砍半","cliffcompaction-truncate-only-agent-compaction","2026-09-23T17:10:56+00:00"]