[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-cliffcompaction-truncate-only-agent-compaction":3,"topics-all":38,"news-related-e4b3903e-c65f-46cb-90ae-81502eb8cdd9":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"e4b3903e-c65f-46cb-90ae-81502eb8cdd9","CliffCompaction开源:只删不改的会话压缩,长程Agent成本砍半","CMU 与博世推出开源代理 CliffCompaction:架在 Claude Code、Codex 前面,超 token 阈值即按「只截断、不改写」压缩会话历史。论文称成本最多降 50%,Terminal-Bench 分数反升,Kimi K2.6 以更低成本追平 Opus 4.7。","长程 coding agent 有个不起眼却持续膨胀的成本:每轮推理都要重发整段会话历史。跑几小时的任务,上下文轻松堆到几十万 token,账单按每轮全量计。CMU 与博世 AI 中心 9 月 22 日放出的 CliffCompaction(arXiv:2609.26779)解法近乎粗暴:在 agent 与 API 之间架一个透明代理,历史一超阈值就压缩——只做两件事:截断和丢弃,绝不改写一个字。\n\n## 反直觉的核心:不总结,只删减\n\n传统 harness(Claude Code、Codex CLI 自带的 auto-compaction)的思路是把旧上下文折叠成一段新写的摘要,问题是「摘要的摘要」会漂移——每压缩一次,事实就离原始内容远一步。CliffCompaction 的规则全部是机械的:系统提示与任务描述原样保留;最近几轮对话原样保留;工具结果不超过 500 字符的整段保留,超长的直接丢弃——理由是文件还在磁盘上,agent 随时能重读;图片在摘要里一律丢弃。阈值默认 20 万 token,每次重新压缩会把上一份压缩结果整个扔掉、只从原始会话重算,论文称之为「never compact a compaction」:漂移不会累积。\n\n实现上,代理对每条消息做哈希形成 hash chain,靠最长前缀匹配替换历史,解析失败一律原样透传(fail-open)。支持 Anthropic 与 OpenAI 双协议,`uv tool install cliffcompaction` 后 `cliff enable` 即用,agent 零改动;shadow 模式只记录会压掉什么,不动真请求。\n\n## 数字说话:成本降一半,分数反升\n\n论文报告:有界上下文下成本最多降 50%,Terminal-Bench 分数维持或提升。论文表格显示,Kimi K2.6 上 SWE-bench Verified(mini-swe-agent)全上下文 73.87%,32K 阈值压缩后 73.27%,几乎无损,压到 8K 才掉到 67.6%;Terminal-Bench 2.0 更反直觉,全上下文 59.16%,32K 压缩后反而升到 61.42%。\n\n更妙的是 test-time scaling 账本:省下的单次 rollout 成本让并行采样变便宜,论文称 Terminal-Bench 能以低于两次完整上下文运行的代价多拿 10 个百分点以上;并行扩展下 Kimi K2.6 追平 Opus 4.7、超过 Opus 4.6 与 GPT-5.3 Codex,成本更低。KernelBench 上它支撑超百万 token 的持续学习,200 步后 CUDA 内核加速 2.23 倍,400 步后 3.58 倍——论文自述超过了专门的搜索算法与训练过的 agent。\n\n## 冷静一点看\n\nREADME 提醒要关掉 scaffold 自带压缩——某些 harness 会原地改写历史,破坏前缀匹配。评测以 coding 场景为主,模型覆盖 Kimi K2.5\u002FK2.6\u002FK2.7 与 GLM 5.1\u002F5 Turbo;「删而不改」在其他场景是否同样无损,待社区验证。项目数天前才开源(MIT),star 还在两位数,离生产级背书尚早。\n\n但对每天烧钱跑长任务的团队,这是零侵入实验:`cliff run --shadow -- claude` 先看它会删什么,再决定要不要省这一半。真正值得记住的洞察是:上下文管理上**保真比聪明更重要**——机械的删除保留了「发生过什么」,聪明的总结每次都在冒险丢掉它。\n\n参考:arXiv:2609.26779 · github.com\u002Fnguyenvuthientrang\u002Fcliffcompaction","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.26779","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"2f746335-06d5-4a25-a728-a5d3ccf7ce56","en","CliffCompaction: truncate-only proxy halves coding-agent costs","CliffCompaction (CMU, MIT): truncate-only proxy compacts agent history. Up to 50% cost cut, Terminal-Bench score rises at 32K, 3.58x KernelBench speedup.","Long-horizon coding agents carry a quietly compounding cost: every turn re-sends the entire session history, so a multi-hour run stacks hundreds of thousands of tokens that get billed again on each request. CliffCompaction (arXiv:2609.26779), released September 22 by Carnegie Mellon University and the Bosch Center for AI, takes an almost blunt approach: a transparent proxy sits between the agent and the API, and once history crosses a token threshold it gets compacted — by truncating or dropping content only, never rewriting a word.\n\n## The counterintuitive core: no summarizing, only cutting\n\nTraditional harnesses like Claude Code and Codex CLI fold older context into freshly written prose, and a summary of a summary drifts away from what actually happened. CliffCompaction's rules are mechanical: system prompts and task descriptions pass through verbatim; the most recent turns are kept whole; tool results under 500 characters are kept verbatim while longer ones are dropped — the files are still on disk, and the agent can re-read them; images are dropped from summaries. The default threshold is 200,000 tokens. Each re-compaction discards the previous compacted output and rebuilds from the live session — \"never compact a compaction\" — so drift cannot accumulate.\n\nOn the engineering side, the proxy canonicalizes and hashes every message into a chain, substitutes compacted history by longest-prefix match, and falls open to verbatim passthrough on any failure. It speaks Anthropic Messages plus OpenAI Chat Completions and Responses, installs with `uv tool install cliffcompaction` plus `cliff enable`, and leaves the agent itself unchanged. A shadow mode (`cliff run --shadow -- claude`) logs what would be compacted without touching real requests.\n\n## The numbers: costs halve, scores rise\n\nThe paper reports cost reductions of up to 50% under a bounded context while Terminal-Bench performance is maintained or improved. Its tables show that on SWE-bench Verified (mini-swe-agent, Kimi K2.6), full context scores 73.87% versus 73.27% at a 32K threshold — nearly lossless — falling to 67.6% only at 8K. On Terminal-Bench 2.0 the score actually goes up under compaction: 59.16% at full context versus 61.42% at 32K.\n\nThe bigger story is test-time scaling economics. Per-rollout savings make parallel sampling cheap: the paper claims over 10 additional percentage points on Terminal-Bench for less than the cost of two full-context runs, and under parallel scaling Kimi K2.6 matches Opus 4.7 while exceeding Opus 4.6 and GPT-5.3 Codex at lower cost. On KernelBench it sustains continual learning over sessions beyond a million tokens, reaching 2.23x CUDA kernel speedups after 200 steps and 3.58x after 400 — which the authors note surpasses specialized search algorithms and trained agents.\n\n## Caveats\n\nThe README warns users to disable the scaffold's own compaction, since some harnesses rewrite history in place and break the prefix matching. Evaluation coverage is coding-centric (SWE-bench, Terminal-Bench, KernelBench) with Kimi K2.5\u002FK2.6\u002FK2.7 and GLM 5.1\u002FGLM 5 Turbo; whether truncate-only compaction is equally lossless in other domains awaits community validation. The repo opened days ago under an MIT license and sits at two-digit stars — hardly production-hardened consensus.\n\nStill, for teams burning money on long agent runs, this is a zero-intrusion experiment: run shadow mode, see what it would cut, then decide whether the ~50% saving is worth it. The insight worth keeping: in context management, fidelity beats cleverness — mechanical deletion preserves what happened, while every clever summary gambles with losing it.\n\nReference: arXiv:2609.26779 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.26779) · GitHub: nguyenvuthientrang\u002Fcliffcompaction","cliffcompaction-truncate-only-agent-compaction","2026-09-23T17:10:56Z","2026-09-23T17:10:59.964356Z","2026-09-23T17:10:59.964367Z",true,"agent",727,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"436fb2b9-4c48-4631-977c-c9539650f975","Kimi K2.7-Code 开源:Moonshot 把\"过度思考\"砍掉三成,长程编程更经济","kimi-k2-7-code-moonshot-30pct-token-cut","2026-06-13T02:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"73bd4795-6651-425b-b407-372d1ec0e793","Ollama 0.24 解锁新玩法：一行命令让 OpenAI Codex 跑在本地开源模型上","ollama-0-24-codex-app-local-open-source","2026-05-27T22:01:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"2dbc7c0f-083a-47d4-ba9a-8a7d66b22ae0","亚马逊八阶段配方:后训练让 GLM-4.5-Air 反超官方版","amazon-rufus-air-post-training-recipe","2026-09-25T23:12:24+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"448dbda7-e4ac-44ba-8a40-7af01012b8df","把仓库\"换脸\"再测:SWE-bench 高分最多掉 14.4 分,编码 agent 疑在背题","schrodinger-repo-swe-bench-memorization","2026-09-24T17:11:54+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"3559e613-9558-48e1-ab20-f53b62796363","让每个 token 用上全部专家:高德 IntBMoE 解耦参与度、计算与显存,60ms 服务数亿用户","intbmoe-full-participation-block-moe","2026-09-21T13:01:54+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"c696208b-6535-4eb9-b1ed-2e4f835d2f88","NVIDIA SoL-Pi 把 coding agent 的 token 砍掉 44%,harness 开始变天","nvidia-sol-pi-harness-token-compression","2026-09-19T03:00:00+00:00"]