Anyone who has run long-horizon agent tasks knows the pattern: every step's reasoning gets stacked into the context verbatim, and before the task finishes, historical reasoning has consumed most of the window. Worse, once an action has executed and the environment has returned feedback, most of that reasoning is never consulted again — yet it still occupies every subsequent request, inflating cache usage, billing, and latency. Static chain-of-thought compression cannot fix this, because agent reasoning is not text frozen at completion: edit the reasoning at step t and the tool call, observation, and decisions at step t+1 all change with it. A 30-page arXiv paper submitted September 24, "When Can Agents Forget Their Reasoning?", confronts the question head-on: when can an agent safely forget reasoning it has already executed?
Entropy as the scalpel, training-free and online
The proposed method is called ICLR (Interaction Aware Compression for Long Horizon Reasoning — yes, it collides with the conference acronym, a pun the paper leans into). The core idea is strikingly simple: after each interaction step, newly generated reasoning is partitioned into blocks and scored by a frozen proxy model — low-entropy blocks are removed, while actions, tool calls, and observations are always preserved. Nothing needs fine-tuning; the loop plugs directly into a real agent interaction cycle, and the compressed history is written back to persistent state, so all later model calls operate on the modified trajectory. The criterion is intuitive: low entropy means the model is confident about this content, which usually marks repeated confirmations and procedural filler, while high-entropy stretches tend to coincide with genuine decision forks.
260 tasks: reward up, tokens down
Across 260 WorkBuddyBench tasks spanning code, office, security, and web domains, with DeepSeek-V4-Flash as the base model, average reward climbs from 0.699 to 0.718 while input tokens fall 25.5%, output tokens 14.4%, and cache-read tokens 33.3%; total tokens drop from 2.69M to 2.19M. Per-task, the method wins 102, ties 76, and loses 82. The domain breakdown is the most interesting part: security jumps 22.9 points, office and web hold roughly flat, and code alone drops 6.4 points — deleting reasoning is not a free lunch, and in domains that demand careful re-reading of history, entropy ranking can still shear off critical context.
Trajectory amplification: delete 10%, save 60%
The paper's most counterintuitive finding is "trajectory amplification": local deletion volume and total system savings are wildly disproportionate. Randomly deleting 10% of reasoning tokens cuts total input by 60.5%; randomly deleting 20% cuts it by only 44.0%. The reason is that any local edit to reasoning history changes subsequent tool selection, observations, and recovery behavior, which then amplifies into a change across the whole trajectory — so the local decision of "how much to delete" relates nonlinearly to the global outcome of "how much is saved". This is why the work insists that agent reasoning compression is a closed-loop decision problem, not a text-editing problem.
The real conclusion: reasoning is working memory, not an archive
Three experiment families close the paper. Representation probing shows the hidden state at the current context boundary encodes a decodable signal of whether a reasoning stretch will be reused later, reaching an AUROC of 0.844. Trajectory analysis yields the more practical rule: when task-critical derived state lives only in reasoning, the future behavioral risk rate is 67.3%; once that state is externalized into code, files, tool outputs, or environmental feedback, the rate falls to 36.2%. Controlled interventions are blunter still: clearing reasoning after relevant code has been written to disk leaves the score unchanged, while hiding observed environmental feedback makes the agent issue another Bash call and observe the same error again. The one-line conclusion: the value of historical reasoning depends on where task-relevant information lives — it is precious while it remains the sole carrier, and becomes disposable working memory once externalized.
The engineering takeaway is practical: rather than agonizing over how much more context window you can buy, first sink agent outputs into files and tool outputs, then let an entropy gatekeeper purge low-entropy reasoning on schedule. The fix for a window that is too small may not be a longer window, but cleaner forgetting. Reference: https://arxiv.org/abs/2609.29875