[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-swe-pruner-pro-bytedance":3,"news-related-6b52b4a9-d567-46b8-99c1-e9c65ba59b16":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"6b52b4a9-d567-46b8-99c1-e9c65ba59b16","SWE-Pruner Pro:ByteDance 让 Agent 自己当剪枝器,省 39% token 还涨分","多轮编程 Agent 跑长任务时,上下文里 80% 是工具返回的「无用行」——测试日志、报错堆栈、diff 残留。SWE-Pruner Pro 反直觉:哪些行该删,Agent 自己的中间层表征已经知道,根本不需要外挂分类器。\\n\\n字节 Seed 这篇 arXiv(2607.18213)在冻结 Coder LLM 上挂小剪枝头,把 Agent 阅读工具输出时的 hidden state 映射成「行级 keep\u002Fprune」决策,叠 length-aware embedding 区分不同长度工具块。在 MiMo-V2-Flash 等两个开源 backbone、四个多轮基准上几乎不增加推理延迟就省下最多 39% 的 prompt+completion token,SWE-Bench Verified 求解率 +3.8pp,Oolong 长上下文准确率 +2.2pp。\\n\\n意义在范式转移。之前所有上下文压缩方法都默认「需要独立判定模型」,Pro 用 Coder 自身表征证明这是多余假设——Agent 已是「上下文相关性」最优判断器,只是没人把这层信号读出来。代码已开源,下一步:能否搬到通用 Agent。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.18213","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"22f928ea-10db-468c-8efc-954ce1e3e6f1","en","SWE-Pruner Pro: agents prune their own context, 39% cheaper","When multi-turn coding Agents run long tasks, 80% of the context is \"useless lines\" from tool returns — test logs, error stacks, diff remnants. SWE-Pruner Pro is counter-intuitive: which lines to delete, the Agent's own middle-layer representations already know — no external classifier is needed. ByteDance Seed's arXiv paper (2607.18213) attaches a small pruning head on top of a frozen Coder LLM, mapping the Agent's hidden state when reading tool output into a \"line-level keep\u002Fprune\" decision, with length-aware embeddings to differentiate tool blocks of different lengths. On two open-source backbones including MiMo-V2-Flash and four multi-turn benchmarks, it saves up to 39% of prompt+completion tokens with almost no added inference latency, lifts SWE-Bench Verified solve rate by +3.8pp, and lifts Oolong long-context accuracy by +2.2pp. The significance is a paradigm shift. Previously, every context-compression method assumed \"an independent decision model is needed\". Pro uses the Coder's own representations to prove that assumption is unnecessary — the Agent is already the optimal \"context-relevance\" judge; it's just that nobody has been reading out that signal. The code is open-sourced; the next question: can it be ported to general Agents.","swe-pruner-pro-bytedance","2026-07-25T12:00:00Z","2026-07-25T12:06:52.414195Z","2026-08-19T02:08:40.142862Z",true,"agent",99,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4bbc55d2-cabc-477f-a3ad-4e2c119aff2a","TokTier 抓住 Agent 推理的隐藏瓶颈：缓存命中 94.1%，分词仍吃掉 64% 首 token 时间","toktier-stateful-tokenization-agent-serving","2026-07-31T17:56:30+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"37aa0bc9-d135-444f-842e-0b40388d29e9","Qwen3.7-Max 原生兼容 Anthropic API 协议：Claude Code 现已可直接调用阿里模型","qwen3-7-max-anthropic-api-claude-code","2026-05-27T10:05:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"57b691cc-476c-4427-8618-e29127654b34","AMD ROCm 7 原生支持 Qwen3-Coder-Next：单卡 256k 上下文打破推理硬件垄断","amd-rocm7-qwen3-coder-next-256k-mono","2026-05-25T16:10:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"0d0e5ce8-fa18-4907-b811-2918ff8464e4","FlexSQL：小型LLM如何在Text-to-SQL任务上超越GPT-o3和DeepSeek-R1","flexsql-nus-text-to-sql-spider2-65pct-gpt-oss-120b","2026-05-05T10:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"69e52a42-19c7-4580-8c49-5446233fbdde","7B模型如何超越GPT-4o？ICLR Oral论文揭示AgentFlow流式训练新范式","agentflow-7b-icrl-oral-flow-grpo-14-9pct","2026-05-03T01:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00"]