Long-horizon coding agents carry a quietly compounding cost: every turn re-sends the entire session history, so a multi-hour run stacks hundreds of thousands of tokens that get billed again on each request. CliffCompaction (arXiv:2609.26779), released September 22 by Carnegie Mellon University and the Bosch Center for AI, takes an almost blunt approach: a transparent proxy sits between the agent and the API, and once history crosses a token threshold it gets compacted — by truncating or dropping content only, never rewriting a word.

The counterintuitive core: no summarizing, only cutting

Traditional harnesses like Claude Code and Codex CLI fold older context into freshly written prose, and a summary of a summary drifts away from what actually happened. CliffCompaction's rules are mechanical: system prompts and task descriptions pass through verbatim; the most recent turns are kept whole; tool results under 500 characters are kept verbatim while longer ones are dropped — the files are still on disk, and the agent can re-read them; images are dropped from summaries. The default threshold is 200,000 tokens. Each re-compaction discards the previous compacted output and rebuilds from the live session — "never compact a compaction" — so drift cannot accumulate.

On the engineering side, the proxy canonicalizes and hashes every message into a chain, substitutes compacted history by longest-prefix match, and falls open to verbatim passthrough on any failure. It speaks Anthropic Messages plus OpenAI Chat Completions and Responses, installs with uv tool install cliffcompaction plus cliff enable, and leaves the agent itself unchanged. A shadow mode (cliff run --shadow -- claude) logs what would be compacted without touching real requests.

The numbers: costs halve, scores rise

The paper reports cost reductions of up to 50% under a bounded context while Terminal-Bench performance is maintained or improved. Its tables show that on SWE-bench Verified (mini-swe-agent, Kimi K2.6), full context scores 73.87% versus 73.27% at a 32K threshold — nearly lossless — falling to 67.6% only at 8K. On Terminal-Bench 2.0 the score actually goes up under compaction: 59.16% at full context versus 61.42% at 32K.

The bigger story is test-time scaling economics. Per-rollout savings make parallel sampling cheap: the paper claims over 10 additional percentage points on Terminal-Bench for less than the cost of two full-context runs, and under parallel scaling Kimi K2.6 matches Opus 4.7 while exceeding Opus 4.6 and GPT-5.3 Codex at lower cost. On KernelBench it sustains continual learning over sessions beyond a million tokens, reaching 2.23x CUDA kernel speedups after 200 steps and 3.58x after 400 — which the authors note surpasses specialized search algorithms and trained agents.

Caveats

The README warns users to disable the scaffold's own compaction, since some harnesses rewrite history in place and break the prefix matching. Evaluation coverage is coding-centric (SWE-bench, Terminal-Bench, KernelBench) with Kimi K2.5/K2.6/K2.7 and GLM 5.1/GLM 5 Turbo; whether truncate-only compaction is equally lossless in other domains awaits community validation. The repo opened days ago under an MIT license and sits at two-digit stars — hardly production-hardened consensus.

Still, for teams burning money on long agent runs, this is a zero-intrusion experiment: run shadow mode, see what it would cut, then decide whether the ~50% saving is worth it. The insight worth keeping: in context management, fidelity beats cleverness — mechanical deletion preserves what happened, while every clever summary gambles with losing it.

Reference: arXiv:2609.26779 (https://arxiv.org/abs/2609.26779) · GitHub: nguyenvuthientrang/cliffcompaction