[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-agent-editing-world-model":3,"topics-all":35,"news-related-237d0bac-204c-4818-9262-e576f96df623":54},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"237d0bac-204c-4818-9262-e576f96df623","Agent 被自己的历史拖累:人大 AEWM 编辑任务状态,六项基准最高涨 6.7 分","长程任务里 Agent 常被历史拖累:未证实假设和过时计划留在上下文,持续扭曲后续决策。人大团队论文 AEWM 把世界模型目标从预测环境观测改为编辑任务状态:先分类决策,再改写噪声延续,直接接入真实执行。自报 70.5% macro-F1 超基线 10.6 分,六项基准平均涨 3.2 至 6.7 分。","跑长程任务的 Agent 都遇到过同一类翻车:任务进行到一半,它突然引用一个从没被证实过的假设,或者照着几轮之前已经过时的计划继续执行。问题往往不在模型能力,而在上下文历史本身——错误一旦写进历史,就会像污染物一样持续扭曲后面的每一步决策。中国人民大学团队 9 月 23 日提交到 arXiv 的论文给这个现象起了名字:任务状态污染(task-state contamination),并给出了一套编辑式的解法。\n\n## 世界模型换了目标:不预测环境,改编辑状态\n\n论文提出的 Agent-Editing World Model(AEWM)首先对主流做法提出质疑:现有语言世界模型通常学习预测环境的下一步观测,但工具返回的结果高熵、依赖具体执行过程,在真实反馈随手可得的情况下,重建这些观测价值有限。AEWM 转而建模推理和动作如何影响未来的任务进展。\n\n具体分两步。Action Judge 负责在执行前给决策分类,区分关键(Critical)、探索(Exploratory)、噪声(Noisy)三类;State Revision 负责从同一观测历史出发,直接改写噪声决策的推理-动作延续。推理框架 EditAct 把两者接进真实执行——它不是在旁边打分给建议,而是直接改变后续决策所依赖的状态。论文还用验证过的 EditAct 轨迹做拒绝采样微调(AEWM-RFT),不需要在线 AEWM 引导,在三个领域比 Self-RFT 高 2.2 到 2.6 分。\n\n## 自报数字:判别器超基线 10.6 分,六基准平均涨 3.2-6.7 分\n\n论文在 Search、Terminal、Software Engineering 三个领域通过 mid-training 加监督微调训练 AEWM。自报成绩单上,AEWM 在其 Action Judge 基准达到 70.5% macro-F1,比最强前沿基线高 10.6 分;EditAct 在六项基准、三种 Agent 骨干上的平均分比最强基线高 3.2 到 6.7 分。\n\n需要说明,这些数字均来自论文作者自己的评测与基准,尚未见第三方独立复现,采信时建议保留归因措辞。\n\n## 对做 Agent 工程的人意味着什么\n\n这条路线的价值在于把纠错时机前移:传统做法是让错误发生、再靠重试或反思挽救;AEWM 则在错误写入历史之前就把噪声决策识别出来并改写。对上下文窗口越拉越长、任务链越来越长的 Agent 系统来说,历史卫生(history hygiene)可能比更大的模型更划算——毕竟污染是随步数累积的,而参数规模并不天然免疫污染。\n\n另一个值得留意的信号:论文署名来自人民大学 RUC 团队(通讯作者 Wayne Xin Zhao、文继荣等),Agent 基础设施这个方向的学术产出正在向中国团队集中。\n\n论文原文见 [arXiv:2609.28416](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28416)。所以呢:下次你的 Agent 在第 20 步引用第 3 步的错误假设时,先别急着换更大的模型——它可能只是需要一次状态编辑。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28416","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"6df53b07-f898-4be1-8501-397be2b8ade0","en","Agents Get a State Editor: AEWM Fixes Contaminated Task History","RUC's AEWM edits agent task state, not environment observations. Authors report 70.5% macro-F1, +10.6 over the best baseline, max +6.7 across six benchmarks.","Long-horizon agents keep tripping over their own history: unverified assumptions and outdated plans persist in context and distort later decisions. A Renmin University of China team names this task-state contamination and proposes AEWM, which reframes the language world model objective from predicting environment observations to editing task state. Action Judge classifies decisions as Critical, Exploratory, or Noisy; State Revision rewrites noisy continuations; EditAct wires both into real execution. Self-reported: 70.5% macro-F1, 10.6 points above the strongest frontier baseline, and 3.2-6.7 point average gains across six benchmarks and three agent backbones.\n\n## A world model that stops predicting the environment\n\nThe paper opens by questioning the mainstream approach: existing language world models learn to predict the next environment observation, yet tool responses are high-entropy and execution-dependent, so reconstructing them adds limited value when real feedback is cheap. AEWM instead models how an agent's reasoning and actions shape future task progress.\n\nThe design has two parts. Action Judge classifies decisions before execution into Critical, Exploratory, and Noisy categories. State Revision rewrites noisy reasoning-action continuations from the same observed history. The inference framework EditAct integrates both with real execution: rather than standing aside and offering critiques, it directly changes the state that subsequent decisions depend on. A further variant, AEWM-RFT, fine-tunes on verified EditAct trajectories via rejection sampling, beating Self-RFT by 2.2-2.6 points across three domains without online AEWM guidance.\n\n## Self-reported numbers: +10.6 on the judge, +3.2-6.7 across six benchmarks\n\nThe team trained AEWM across Search, Terminal, and Software Engineering through mid-training plus supervised fine-tuning. On their own Action Judge benchmark, AEWM reaches 70.5% macro-F1, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2 to 6.7 points over the strongest baseline.\n\nOne caveat: all figures come from the authors' own evaluation; no independent third-party replication is out yet, so treat them as self-reported.\n\n## Why this matters for agent engineering\n\nThe core value is moving the correction point earlier. The traditional loop lets errors happen, then retries or reflects; AEWM identifies and rewrites noisy decisions before they contaminate the history. As context windows stretch and task chains grow longer, history hygiene may prove cheaper than larger models: contamination accumulates with steps, while parameter count buys no immunity.\n\nA signal worth noting: the author list comes from Renmin University's RUC team (corresponding authors Wayne Xin Zhao and Ji-Rong Wen, among others), and academic output on agent infrastructure is visibly concentrating in Chinese teams.\n\nFull paper at [arXiv:2609.28416](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28416). The takeaway: next time your agent cites a step-3 wrong assumption at step 20, don't rush to swap in a bigger model. It may just need a state edit.","agent-editing-world-model","2026-09-29T17:08:51Z","2026-09-29T17:10:02.651364Z","2026-09-29T17:10:02.651379Z",true,"agent",127,[36,45],{"slug":37,"tag_slug":37,"title_zh":38,"title_en":39,"intro_zh":40,"intro_en":41,"id":42,"is_active":32,"created_at":43,"modified_at":44},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":32,"created_at":52,"modified_at":53},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":55},[56,61,66,71,76,81],{"id":57,"title":58,"news_slug":59,"published_at":60},"fe5546f3-09e4-42a3-bfde-6bbd0e8d474c","外星世界实测:探索 4 轮 87.6%,垫底 12.9%","explorationbench-alien-worlds","2026-09-28T21:07:30+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"dcd8b3e1-a3c7-4614-aba4-9002219ea5f6","LibreDB Studio 0.15 发布:本地 LLM 接管数据库交互","libredb-studio-local-llm-agent","2026-09-15T00:00:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"31c09fea-8993-4f10-b683-499672fcafe3","世界模型不能再靠爬视频硬堆:游戏引擎补上了缺失的奖励信号","game-engine-rlhev-world-models","2026-08-30T13:10:00+00:00",{"id":72,"title":73,"news_slug":74,"published_at":75},"5a90a793-8ec1-4b3a-9691-edef5ffe8535","AI「思想病毒」实证:Anthropic 与 EPFL 让恶意想法在 Agent 间自我复制,免疫只需一段警告","mind-viruses-multi-agent-llm","2026-08-18T13:30:00+00:00",{"id":77,"title":78,"news_slug":79,"published_at":80},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":82,"title":83,"news_slug":84,"published_at":85},"3c6fcf46-f5bb-4136-931c-69cd64216e12","Skill-Use 基准揭示 Agent 短板：会做任务，不等于会用 Skill","skill-use-agent-harness-benchmark","2026-08-06T08:00:00+00:00"]