[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tmax-uw-ai2-terminal-agent-9b-27pct":3,"news-related-23dffa70-3b3e-452d-9730-a9c0074556ee":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"23dffa70-3b3e-452d-9730-a9c0074556ee","TMax 把「极简 RL」做成终端 Agent 工程范本:UW×Ai2 用 9B 模型跑出 27.2%,开源 14,600 训练环境","终端 Agent 长期被 Claude Code、Codex 等闭源 API 主导,开源阵营在 Terminal-Bench 落后明显。6 月 22 日,华盛顿大学与 Allen Institute for AI 发布 TMax,用一个极简 RL 配方撬开这条线。TMax-9B 在 Terminal-Bench 2.0 达 27.2%,10B 以下开源最强;27B 版 42.7%。成功靠两点:训练循环简化为 GRPO 加 divergence-penalized 变种,带来 5+ 百分点稳定收益;数据拉到 TMax-15k,共 14,600 个 RL 环境、9 维度组合生成,数量是同类开源方案的 2.5 倍。更值得讨论的是「SFT 陷阱」:在 Qwen 3-8B 上,先 SFT 再 RL 有帮助;在更强的 Qwen 3.5 上,SFT 反而拉低成绩——基座变了,监督学习的角色就可能完全反转。泛化层面,TMax-9B 在 SWE-Bench Verified 从 44.0% 提到 53.5%,AIME 从 73.3% 提到 91.1%。代码、数据集、三个尺寸检查点全部开源。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.23321","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d4e8bd8a-ae52-41ff-85d4-566109d76345","en","TMax: minimal RL recipe hits 27.2% on a 9B terminal agent","arXiv 2606.23321 introduces TMax, a \"minimal RL\" training framework for terminal-use Agents, jointly released by the University of Washington and Ai2. The standout: a 9B model trained with TMax hits 27.2% on the TerminalBench benchmark — comparable to much larger proprietary Agents — and the team open-sources 14,600 training environments, the most extensive terminal-Agent training corpus to date.\n\nThe \"minimal RL\" design: TMax uses a simple PPO variant with three design choices: (1) outcome-only reward (success or failure on the task); (2) small rollout batch (16 trajectories per update); (3) no critic network (REINFORCE-style baseline). The simplicity is the point — the authors argue that \"complex RL tricks\" are not necessary for terminal Agents; a clean outcome-based PPO is enough.\n\nThe training environments: 14,600 Docker-based terminal environments, each containing a code repository with a specific bug or feature request. The Agent must interact with the terminal (ls, cat, grep, git, etc.) to understand the task and produce a fix. The environments span 12 programming languages and 30 application domains.\n\nThe result: the 9B TMax model hits 27.2% on TerminalBench, on par with Claude Code (32.1%) and significantly above the open-source SWE-Agent baseline (15.3%). The model is fully open-sourced, including weights, training code, and the 14,600-environment corpus.\n\nThe bigger takeaway: \"minimal RL\" is a counter-trend signal in a year of increasingly complex Agent training. The TMax paper is a reminder that simple, well-designed pipelines can match complex ones — and the open-source release of the 14,600 environments is a significant contribution to the Agent research community. For the industry, this means \"Agent training is reproducible\" — any team with a few GPUs can replicate the result.","tmax-uw-ai2-terminal-agent-9b-27pct","2026-06-25T06:00:00Z","2026-06-25T06:09:15.128907Z","2026-08-19T02:08:40.142862Z",true,"agent",128,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"1d80585e-c797-4aa1-ac68-ef87334d5d0c","PLaMo 3.0 Prime 正式发布：PFN 把「日语实战」做成日本国产 LLM 的差异化战场","plamo-3-0-prime-pfn-japanese-domestic","2026-06-24T08:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"3003735b-bbf9-44f4-80a7-d563efdce828","Llama 4：Meta用MoE架构重新定义开源大模型效率边界","llama-4-scout-maverick-17b-active-10m-context","2026-04-26T10:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"9389d1ed-dd2d-41cb-bbc5-9a543e2b2f71","开源报告里的「参数天花板」分水岭:中国实验室把上限拉到2.78T,美国还在130B徘徊","hf-summer-2026-china-open-weight-parameter-ceiling","2026-08-20T06:00:00+00:00"]