[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-agents-forget-reasoning-iclr-compression":3,"topics-all":38,"news-related-5fa17d26-6090-4805-893a-aa880cf33369":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"5fa17d26-6090-4805-893a-aa880cf33369","让 Agent 学会遗忘:删掉旧推理,分数反而涨了","长程 Agent 把历史推理背在上下文,token 越滚越大。论文提出免训练在线压缩方法 ICLR:冻结代理模型按熵值排序推理块,低熵即删,动作与观察全保留。260 个 WorkBuddyBench 任务平均奖励从 0.699 升至 0.718,输入省 25.5%、缓存读省 33.3%。推理外化进代码文件后即可丢弃。","跑过长时间任务的 Agent 都会遇到同一件事:每一步的思考过程被原封不动地堆进上下文,任务还没做完,历史推理已经吃掉了大半窗口。更麻烦的是,这些推理在对应的动作执行完、环境反馈也拿到之后,大部分就再没被用过——但它们依然占据着每一轮请求的输入,缓存、计费、延迟全都跟着涨。静态的思维链压缩解决不了这个问题,因为 Agent 的推理不是写完就定的文本:改掉第 t 步的推理,第 t+1 步的工具调用、观察和后续决策全都会跟着变。arXiv 9 月 24 日的 30 页论文《When Can Agents Forget Their Reasoning?》正面回答:什么时候,Agent 能安全忘掉已执行过的推理。\n\n## 用熵当手术刀,免训练在线删\n\n论文提出的方法叫 ICLR(Interaction Aware Compression for Long Horizon Reasoning,和那个顶会缩写撞名,算论文自带的梗),核心思路出奇地朴素:在每一轮交互结束后,把新产生的推理切块,交给一个冻结的代理模型打分——低熵的块删掉,动作、工具调用和环境观察一律保留。全程无需微调,直接嵌进真实的 Agent 交互循环,压缩后的历史写回持久层,后续所有模型调用都在改过的轨迹上继续跑。判断标准很直觉:低熵意味着模型对这些内容很有把握,大概率是重复确认和程序性叙述;高熵的部分往往对应真正的分叉决策。\n\n## 260 个任务:奖励涨了,token 降了\n\n在 260 个 WorkBuddyBench 任务(覆盖代码、办公、安全、网页四个域)上,底座模型为 DeepSeek-V4-Flash:平均奖励从 0.699 升到 0.718,同时输入 token 减少 25.5%,输出 token 减少 14.4%,缓存读 token 减少 33.3%,总 token 从 2.69M 降到 2.19M。逐任务看是 102 胜 76 平 82 负。最有意思的是分域表现:安全域大涨 22.9 分,办公域和网页域基本持平,唯独代码域掉了 6.4 分——删推理不是免费的午餐,在需要精读历史的场景里,熵值排序也会误伤关键上下文。\n\n## 轨迹放大:删 10%,省下 60%\n\n论文里最反直觉的发现叫「轨迹放大」(trajectory amplification):局部删除量和系统总算力节省完全不成比例。随机删 10% 的推理 token,总输入降了 60.5%;随机删 20%,反而只降 44.0%。原因在于,对推理历史的任何局部改动都会改变后续的工具选择、观察内容和恢复行为,进而放大成整条轨迹的变化——「删多少」和「省多少」之间是非线性的。这也说明 Agent 推理压缩是闭环决策问题,不是文本编辑问题。\n\n## 真正的结论:推理是工作记忆,不是档案\n\n论文用三组实验收尾。表征探针显示,当前上下文边界的隐状态里确实编码了「这段推理将来还要不要用」的信号,AUROC 达到 0.844。轨迹分析给出了更实用的规律:当任务关键状态只存在于推理里时,未来行为风险率是 67.3%;一旦这些状态被外化成代码、文件、工具输出或环境反馈,风险率降到 36.2%。受控干预更直接:代码写进磁盘之后再清掉后续推理,得分纹丝不动;而把环境反馈藏起来,Agent 会多发一次 Bash 调用、重新观察同一个报错。结论一句话:历史推理的价值取决于「任务相关信息存在哪里」——它还是唯一载体时很值钱,外化之后就变成可丢弃的工作记忆。\n\n对工程团队的启示:与其纠结上下文窗口还能再买多大,不如先把 Agent 的产出沉淀到文件和工具输出里,再配一个熵值门卫按时清理低熵推理。窗口不够用的解法,可能不是更长的窗口,而是更干净的遗忘。参考:https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.29875","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.29875","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"5c6dea30-5b29-456f-bce8-145445b757e5","en","Agents Can Forget: Deleting Old Reasoning Helps","Training-free ICLR ranks reasoning blocks by entropy and deletes low-entropy ones. Reward rises to 0.718 on 260 tasks, input tokens fall 25.5%.","Anyone who has run long-horizon agent tasks knows the pattern: every step's reasoning gets stacked into the context verbatim, and before the task finishes, historical reasoning has consumed most of the window. Worse, once an action has executed and the environment has returned feedback, most of that reasoning is never consulted again — yet it still occupies every subsequent request, inflating cache usage, billing, and latency. Static chain-of-thought compression cannot fix this, because agent reasoning is not text frozen at completion: edit the reasoning at step t and the tool call, observation, and decisions at step t+1 all change with it. A 30-page arXiv paper submitted September 24, \"When Can Agents Forget Their Reasoning?\", confronts the question head-on: when can an agent safely forget reasoning it has already executed?\n\n## Entropy as the scalpel, training-free and online\n\nThe proposed method is called ICLR (Interaction Aware Compression for Long Horizon Reasoning — yes, it collides with the conference acronym, a pun the paper leans into). The core idea is strikingly simple: after each interaction step, newly generated reasoning is partitioned into blocks and scored by a frozen proxy model — low-entropy blocks are removed, while actions, tool calls, and observations are always preserved. Nothing needs fine-tuning; the loop plugs directly into a real agent interaction cycle, and the compressed history is written back to persistent state, so all later model calls operate on the modified trajectory. The criterion is intuitive: low entropy means the model is confident about this content, which usually marks repeated confirmations and procedural filler, while high-entropy stretches tend to coincide with genuine decision forks.\n\n## 260 tasks: reward up, tokens down\n\nAcross 260 WorkBuddyBench tasks spanning code, office, security, and web domains, with DeepSeek-V4-Flash as the base model, average reward climbs from 0.699 to 0.718 while input tokens fall 25.5%, output tokens 14.4%, and cache-read tokens 33.3%; total tokens drop from 2.69M to 2.19M. Per-task, the method wins 102, ties 76, and loses 82. The domain breakdown is the most interesting part: security jumps 22.9 points, office and web hold roughly flat, and code alone drops 6.4 points — deleting reasoning is not a free lunch, and in domains that demand careful re-reading of history, entropy ranking can still shear off critical context.\n\n## Trajectory amplification: delete 10%, save 60%\n\nThe paper's most counterintuitive finding is \"trajectory amplification\": local deletion volume and total system savings are wildly disproportionate. Randomly deleting 10% of reasoning tokens cuts total input by 60.5%; randomly deleting 20% cuts it by only 44.0%. The reason is that any local edit to reasoning history changes subsequent tool selection, observations, and recovery behavior, which then amplifies into a change across the whole trajectory — so the local decision of \"how much to delete\" relates nonlinearly to the global outcome of \"how much is saved\". This is why the work insists that agent reasoning compression is a closed-loop decision problem, not a text-editing problem.\n\n## The real conclusion: reasoning is working memory, not an archive\n\nThree experiment families close the paper. Representation probing shows the hidden state at the current context boundary encodes a decodable signal of whether a reasoning stretch will be reused later, reaching an AUROC of 0.844. Trajectory analysis yields the more practical rule: when task-critical derived state lives only in reasoning, the future behavioral risk rate is 67.3%; once that state is externalized into code, files, tool outputs, or environmental feedback, the rate falls to 36.2%. Controlled interventions are blunter still: clearing reasoning after relevant code has been written to disk leaves the score unchanged, while hiding observed environmental feedback makes the agent issue another Bash call and observe the same error again. The one-line conclusion: the value of historical reasoning depends on where task-relevant information lives — it is precious while it remains the sole carrier, and becomes disposable working memory once externalized.\n\nThe engineering takeaway is practical: rather than agonizing over how much more context window you can buy, first sink agent outputs into files and tool outputs, then let an entropy gatekeeper purge low-entropy reasoning on schedule. The fix for a window that is too small may not be a longer window, but cleaner forgetting. Reference: https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.29875","agents-forget-reasoning-iclr-compression","2026-09-28T19:30:00Z","2026-09-28T19:09:52.668918Z","2026-09-28T19:09:52.668933Z",true,"agent",435,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"0bb9aed0-f4f6-4d1a-974a-88e47010815e","长上下文压成答案导向记忆:CMC 让冻结 LLM 省一半显存","cmc-context-memory-embedding-long-context-compression","2026-10-07T00:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"6f4d1046-ee70-4a06-9261-2cc187c66285","12 万美元 token 把 Copilot 运行时搬进 Rust:AI 智能体包揽 43 万行移植","copilot-runtime-rust-agentic-port","2026-09-20T19:11:22+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"0565190a-0bcd-492f-934f-0ad2ab32f485","70万参数2.8MB填一张表:Cua开源CUA-S1,单次前向替代23轮LLM","cua-s1-forms-system-one-model","2026-09-20T13:11:48+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"13378d5e-2440-496d-8c3c-7d36858e641d","不聊天的端侧基座:Needle 3 用 8-29MB 在微控制器上跑工具调用","needle-3-tiny-tool-calling-model","2026-09-19T13:09:46+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"ddb7bc6c-6b6e-4797-ab76-d1aeab5a3002","压缩得好≠部署得好:树莓派实测边缘 LLM,LoRA恢复模型100题押97个同答案","edge-llm-compression-raspberry-pi","2026-08-23T13:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"6b203495-fcab-4afe-baa7-1079cf993796","拆开 GLM-5.3 的「后训练工厂」:基座一字未动,靠环境合成与 1e-7 对齐撑起全部提升","glm-5-3-post-training-stack-deep-dive","2026-08-17T13:00:00+00:00"]