[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-compactionrl-glm-5-2":3,"news-related-49d0aaf8-6fcf-4bf9-83fb-19a016ae2784":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"49d0aaf8-6fcf-4bf9-83fb-19a016ae2784","CompactionRL:把上下文压缩塞进 RL 循环,GLM-5.2 训练管线吃下 5–7pp 编码代理增益","智谱 (Z.ai) 团队 7 月 6 日在 arXiv 公开 CompactionRL (2607.05378),把「上下文压缩」从推理期招数重写为可训练的 RL 原语,直接服务 GLM-5.2 (750B-A40B) 的训练管线。方法用跨轨迹 GAE + token 级 loss 归一化,把任务执行与摘要生成联合优化,让模型在同一份策略里既会干活又会压缩。\n\n效果在开源 MoE 上复现得很整齐:GLM-4.5-Air (106B-A30B) SWE-bench Verified Pass@1 拉到 66.8% (+7.0)、Terminal-Bench 2.0 24.5% (+3.1);GLM-4.7-Flash (30B-A3B) 同两项分别 +5.5、+6.8,跑出 56.0% \u002F 20.2%。\n\n更值得注意的是其解决的真问题:长程 agent 轨迹一旦被压缩,GRPO 这类基于 group-level advantage 的方法会失效,因为同 prompt 不同 rollout 的子轨迹数与长度不再对齐;CompactionRL 切到 critic-based PPO + token 级优势估计,把 RL 的训练假设从「固定长度轨迹」改成「可变长子轨迹」。\n\n工程上也给出基座:训练\u002Frollout 用智谱自研 slime 框架,支持 white\u002Fblack-box rollout、sub-agent 工作流与 FP8 KV-cache,正好承载 GLM-5.2 的并行 OPD 后训练负载。这意味着 RL 流水线首次把「上下文预算」当成可学习的旋钮,而不只是推理时的兜底——对所有还在纠结 agent 长程规划的团队,这是一份可复用的训练配方。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.05378","1eab5c4a-0c8e-49a4-8ac8-0f84a2a3c3a4",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"77012ebf-1166-47be-9d7b-4b35f00c4b77","en","CompactionRL: context compression in the RL loop, +5-7pp coding gains","The Zhipu (Z.ai) team publicly released CompactionRL (2607.05378) on arXiv on July 6, rewriting \"context compression\" from an inference-time trick into a trainable RL primitive, directly serving GLM-5.2's (750B-A40B) training pipeline. The method uses cross-trajectory GAE + token-level loss normalization, jointly optimizing task execution and summary generation, letting the model both work and compress in the same policy. The effect reproduces neatly on open-source MoE: GLM-4.5-Air (106B-A30B) SWE-bench Verified Pass@1 pulled to 66.8% (+7.0), Terminal-Bench 2.0 24.5% (+3.1); GLM-4.7-Flash (30B-A3B) on the same two items gets +5.5, +6.8, running 56.0% \u002F 20.2%. More noteworthy is the real problem it solves: once long-horizon agent trajectories are compressed, methods like GRPO based on group-level advantage will fail, because the sub-trajectory count and length of different rollouts for the same prompt no longer align; CompactionRL switches to critic-based PPO + token-level advantage estimation, changing the RL training assumption from \"fixed-length trajectories\" to \"variable-length sub-trajectories\". The engineering side also provides a base: training\u002Frollout uses Zhipu's self-developed slime framework, supporting white\u002Fblack-box rollout, sub-agent workflows, and FP8 KV-cache, perfectly carrying GLM-5.2's parallel OPD post-training workload. This means the RL pipeline has for the first time treated \"context budget\" as a learnable knob rather than just an inference-time fallback — for all teams still struggling with agent long-horizon planning, this is a reusable training recipe.","compactionrl-glm-5-2","2026-07-10T12:08:00Z","2026-07-10T12:08:20.972757Z","2026-08-19T02:08:40.142862Z",true,"agent",369,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"fa8a0c53-dd89-4870-8e6e-978ce038a919","GitHub Copilot 上线 Kimi K2.7：开源权重模型首次进入默认模型选择器","github-copilot-kimi-k2-7","2026-07-03T06:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"61222f5d-b7e6-4200-8517-3b1972040d24","小米开源 MiMo Code：把 Coding Agent 拆成「计算-记忆-演化」三段式，长程编程首次跑通","xiaomi-mimo-code-compute-memory-evolution","2026-06-11T04:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"fc0cd2d4-cc5f-4f97-b91f-08719e41e8ec","Qwen3.6-27B：27B密集模型超越397B MoE，单卡部署的编程新选择","qwen-3-6-27b-dense-beats-397b-moe-coding-77pct","2026-04-24T03:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"d4d40e4b-04e9-45cf-a04c-792aca45b152","Kimi K2.6开源发布：万亿参数MoE模型的长时编程与Agent Swarm突破","kimi-k2-6-trillion-moe-1t-32b-active-256k","2026-04-24T03:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00"]