[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-neohorse-1-agentic-post-training-rsi":3,"topics-all":38,"news-related-44aa8908-1cef-48d6-b224-7de12a8d4afd":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"44aa8908-1cef-48d6-b224-7de12a8d4afd","NeoHorse-1：让 Agent 执行轨迹进入自我改进回路","NeoHorse-1 用智能路由、结构化反馈和课程式蒸馏，把 Agent 的执行轨迹变成下一轮训练数据。论文公布 4B 与 9B 两个版本：十项基准的平均分分别从 Qwen3.5 基线的 58.94 提升到 64.87、从 65.60 提升到 69.04，并以 Apache 2.0 开源模型和代码。","Agent 的训练数据，正在从“标准答案”变成“执行现场”。NeoHorse-1 的思路不是再堆一批静态问答，而是把模型写代码、调用工具、接收环境反馈、完成或失败的完整轨迹留下来，再用这些记录决定下一轮学什么。\n\n## 先看结果：小模型也能吃到训练红利\n\nTokenRhythm 发布的 NeoHorse-1 有 4B 和 9B 两个版本，基座分别是 Qwen3.5-4B 与 Qwen3.5-9B。论文在十项基准上评估 Agent 执行、工具调用、代码生成和指令遵循。4B 版本的平均分从同尺寸 Qwen3.5-4B 的 58.94 提升到 64.87；9B 版本则从 65.60 提升到 69.04。关键不是某个榜单抬头，而是提升分布在多个交互任务上。\n\n不过，表格也没有把一切都说成单向胜利。9B 版本在 IFEval 上是 89.09，低于 Qwen3.5-9B 的 89.46；LiveCodeBench v6 两者都是 65.14。后训练更明显地改善了持续操作、工具交互和失败恢复，对静态指令遵循的增益并不整齐。\n\n## 它到底训练了什么\n\n普通指令微调把输入和答案配成一对，但 Agent 的真实工作不是这样结束的。一条轨迹里还有工具调用、工具返回、环境观察、重试动作和最终结果。NeoHorse-1 将这些执行记录组织成 user turn 和 subscene，保留当前轮里的推理、工具调用与可见回复，让模型学习的不只是“应该说什么”，还有“在什么状态下做什么动作”。\n\n第二步是路由。系统背后是异构模型池。路由器会记录任务需要什么能力、实际选择了哪一档服务，以及后续交互如何展开。用户覆盖、服务可用性和部署策略都会改变路由结果，实际调用的模型身份不等同于任务难度。NeoHorse-1 更看重路由对能力需求的估计，把样本安排成三阶段课程：逐步加入更高需求的交互，同时保留一部分低需求样本。\n\n这还不够。离线轨迹里的回复是已经走过的路径，部署时学生模型却会生成自己的前缀。为缩小差距，NeoHorse-1 加入 routing-guided on-policy distillation：学生从记录过的上下文出发生成回复，教师在学生真正走到的前缀上提供监督。蒸馏不只模仿教师的完整答案，也直面学生自己暴露的行为状态。\n\n## “递归自我改进”目前还只是一个闭环原型\n\n论文值得关注的地方不是把 RSI 说得多宏大，而是把它拆成能检查的循环：路由和 Agent 执行产生轨迹；轨迹经过结构校验、六维语义评估和 subscene 标注；能力反馈再分配下一轮训练数据；更新后的模型回到 harness 继续执行。\n\n“自我改进”不再只是让模型自己写一份更好的答案，而是让系统利用真实执行留下的证据，判断下一轮训练补什么。但问题也清楚：目前公开结果证明的是一次后训练有效，不等于模型能无限递归变强。跨多轮迭代是论文的下一步。\n\n代码和模型已公开：4B、9B 权重以 Apache 2.0 发布，仓库提供 BF16 与 8-bit、5-bit、4-bit 量化版本；原生上下文 262,144 tokens，可扩展到 1,010,000 tokens。部署示例覆盖 SGLang 和 vLLM，实际容量取决于显存与服务配置。\n\n我的判断：NeoHorse-1 的价值不在“会自己升级”的神话，而在它把 Agent 训练中最易被浪费的东西——失败、重试、工具返回和路由选择——变成可回收的训练信号。下一轮模型竞争，拼的不是静态答案漂亮，而是谁能把每次执行留下的证据用得更彻底。\n\n资料： [arXiv:2609.08183](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.08183)；[GitHub](https:\u002F\u002Fgithub.com\u002FTokenRhythm\u002FNeoHorse)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.08183","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"787c70e1-d7fd-4bfe-beca-2135a5e7d753","en","NeoHorse-1: Agent Traces Enter a Self-Improvement Loop","NeoHorse-1 turns agent traces into training signals via routing-guided distillation; 4B and 9B models lift ten-benchmark averages to 64.87 and 69.04.","Agent training data is moving from “standard answers” toward the execution scene itself. NeoHorse-1 does not simply add another collection of static question-answer pairs. It records how a model writes code, calls tools, receives environmental feedback, retries, and eventually succeeds or fails. Those records are then used to decide what the next training round should learn.\n\n## The result: smaller models benefit from targeted post-training\n\nTokenRhythm’s NeoHorse-1 comes in 4B and 9B versions, based on Qwen3.5-4B and Qwen3.5-9B. The paper evaluates agent execution, tool use, coding, and instruction following across ten benchmarks. The 4B model’s average rises from 58.94 for the same-size Qwen3.5-4B baseline to 64.87. The 9B model rises from 65.60 to 69.04. The important point is not a sudden jump on one leaderboard, but gains spread across several interactive tasks.\n\nThe table also does not show a uniformly positive story. On IFEval, the 9B model scores 89.09, below Qwen3.5-9B at 89.46. On LiveCodeBench v6, both score 65.14. Post-training is more visible on sustained operation, tool interaction, and failure recovery; gains on static instruction following are not uniform.\n\n## What is actually being trained\n\nA conventional instruction-tuning example often pairs an input with an answer. An agent does not work that way. Its trajectory also contains tool calls, tool returns, environmental observations, retries, and an outcome. NeoHorse-1 organizes these records into user turns and subscenes, retaining reasoning, tool calls, and visible responses in the current turn. The model therefore learns not only what to say, but what action to take in a particular state.\n\nThe second piece is routing. The system is backed by a heterogeneous model pool rather than a single isolated model. The router records the estimated capability demand, the selected service tier, and the interaction that follows. The authors stress that the identity of the model actually served is not a simple proxy for task difficulty: user overrides, service availability, and deployment policy can all affect routing. NeoHorse-1 instead uses the router’s estimate of capability demand to arrange a three-stage curriculum. Lower-demand examples enter progressively, higher-demand interactions are added later, and some lower-demand coverage is retained.\n\nThat still leaves a mismatch. An offline trajectory contains a response that has already been produced, while deployment begins with prefixes generated by the student itself. NeoHorse-1 addresses this with routing-guided on-policy distillation. The student generates a response from a recorded context, and a teacher supervises the prefixes the student actually visits. Distillation therefore also responds to behavioral states exposed by the student.\n\n## Recursive self-improvement is currently a feedback-loop prototype\n\nThe strongest part of the paper is not its grand wording around RSI, but its decomposition into a checkable loop: routing and agent execution produce trajectories; the trajectories pass structural validation, six-dimensional semantic evaluation, and subscene labeling; capability feedback reallocates the next training mixture; the updated model returns to the harness.\n\n“Self-improvement” here means using evidence from real execution to identify what the next training round should cover. The limitation is equally clear. The results demonstrate that one post-training cycle is effective; they do not demonstrate that a model can recursively become stronger without bound. The paper lists multi-iteration extension and broader task settings as next steps, so stability remains an open question.\n\nThe project releases 4B and 9B weights under Apache 2.0, together with code. The repository also provides BF16 and 8-bit, 5-bit, and 4-bit quantized versions. The native context length is 262,144 tokens, while the repository says the base capability can be extended up to 1,010,000 tokens. Deployment examples cover SGLang and vLLM, although practical capacity still depends on GPU memory and serving configuration.\n\nMy view is that NeoHorse-1 matters less because it has already fulfilled the fantasy of a model that upgrades itself, and more because it recycles what agent training usually wastes: failures, retries, tool returns, and routing decisions. The next model race may not be decided by whose static answer looks better, but by whose system uses the evidence left by every execution more completely.\n\nSources: [arXiv:2609.08183](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.08183); [NeoHorse GitHub](https:\u002F\u002Fgithub.com\u002FTokenRhythm\u002FNeoHorse)","neohorse-1-agentic-post-training-rsi","2026-09-09T07:19:25Z","2026-09-09T07:19:40.473709Z","2026-09-09T07:19:40.473723Z",true,"agent",157,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"2fb1e49d-6021-4b40-9032-57ccbf95015e","腾讯混元 Hy3 正式版开源:295B MoE 锁定「Agent 实用性」主战场","tencent-hy3-official-launch","2026-07-06T08:05:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"dcd8b3e1-a3c7-4614-aba4-9002219ea5f6","LibreDB Studio 0.15 发布:本地 LLM 接管数据库交互","libredb-studio-local-llm-agent","2026-09-15T00:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"089f56f3-32ff-4036-89b5-728d5f5a9359","边聊边干活:腾讯混元开源全模态交互 Agent Gander,小脑管对话、大脑管执行","hunyuan-gander-omni-interaction-agent","2026-09-09T21:07:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"9822a1a7-0014-4bd5-bbe0-492401fe6b96","AllSpark 把搜索 Agent 推到 BrowseComp 88.6:SFT-RL Climbing 与推理时上下文管理","allspark-iris-search-agent-sft-rl-climbing","2026-09-07T07:11:17+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"63c30bcd-3ffc-47c5-bd74-c2a9ed8f7c94","DeepSeek Harness 预览版开源:Agent 被拆成可插拔的插件栈,模型只负责想、Harness 负责做事","deepseek-harness-plugin-stack","2026-09-05T06:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"089195fb-7fe5-4ba9-a4bc-8e356fe5e923","BAAI把1000个GitHub仓库蒸馏成5000个技能,科研agent奖牌率31%冲到73%","baai-disco-repo-to-skill-library","2026-09-03T17:07:35+00:00"]