[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-epiphenomenal-cot-55pct-useless-thinking":3,"news-related-e4c13922-29e1-41a9-9470-8dae80f62368":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e4c13922-29e1-41a9-9470-8dae80f62368","推理模型的「无效思考」:55% 的 CoT 步骤对答案概率毫无影响","Chain-of-Thought（思维链）已经成为推理时计算 scaling 的主流范式，DeepMind 的 Co-Scientist、OpenAI 的 o 系列、智谱的 GLM-Z 推理增强版本，几乎都在堆叠 CoT 的长度来换精度。但 arXiv 最新论文 **\"Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models\"**（arXiv:2606.13603）撕开了这层表象——大量看起来在「深思熟虑」的推理步骤，其实对最终答案概率毫无影响。\n\n作者用 early-exit 估计每一步 CoT 的因果重要性，发现推理过程存在一个 sharp 的 **commitment boundary（承诺边界）**：模型从瞬态的中间猜测，突然切换到稳定的高置信答案，而这个切换往往只发生在**一个步骤**之内，远早于推理块结束。边界之后跟着的，是大段 **epiphenomenal（副现象）CoT**——继续写、继续解释、继续「自我核对」，但最终 token 的概率分布几乎纹丝不动。换句话说，模型在「自言自语」。\n\n更妙的是，作者用 attention probe 就能从中间隐藏态线性解码出「答案已经形成」这件事，并且这个 probe 能稳健泛化到未见过的任务。基于这个信号做 early-exit，**CoT 长度平均可以砍掉 55%，性能几乎不掉**。这意味着推理模型的 token 预算至少有一半是浪费的。\n\n这件事的意义不只是省钱。它改变了我们理解推理模型的范式：长 CoT ≠ 真正在思考，而是「先想出答案，再写一段听起来像思考的解释」。对 R1\u002Fo3 类模型的 RLHF 训练目标、早退调度策略、以及 agent 的工具调用规划，都给出了非常具体的优化方向——下次有人告诉你「这个推理模型能写 3000 token 的思考过程」，可以先问一下：那 3000 token 里到底有多少是 commitment boundary 之后的「空转」？","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.13603","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"05e04414-5228-43dc-9571-35c2a9ffab75","en","55% of chain-of-thought steps change nothing about the answer","arXiv 2606.13603 investigates a striking phenomenon in reasoning models: 55% of the CoT (Chain-of-Thought) steps in a typical reasoning chain have zero effect on the final answer probability. The \"useless thinking\" is a significant source of inefficiency, and the paper proposes methods to identify and eliminate it.\n\nThe \"useless thinking\" measurement: the authors use a \"counterfactual\" analysis — for each CoT step, they remove the step and measure the change in the final answer probability. If the probability doesn't change, the step is \"useless.\" Across a set of 1,000 reasoning chains, 55% of steps are useless on average.\n\nThe \"where are the useless steps\" analysis: useless steps are concentrated in the \"exploration\" phase of reasoning (the model tries multiple approaches, most of which are abandoned) and the \"verification\" phase (the model re-checks its work, often redundantly). The \"execution\" phase (the model commits to an approach and works through it) has very few useless steps.\n\nThe fix: a \"step pruning\" method that removes useless steps during inference. The pruning is done in real time, with a small classifier predicting whether each step is useless. The pruned reasoning chains are 50% shorter, with no quality loss. The latency is also reduced, since fewer tokens need to be generated.\n\nThe bigger takeaway: \"reasoning efficiency\" is a real engineering discipline. The \"more thinking is better\" assumption is breaking, and the \"prune the useless thinking\" approach is significantly more efficient. For the industry, this means reasoning models will adopt step-pruning techniques, and the next round of reasoning efficiency improvements will come from this direction.","epiphenomenal-cot-55pct-useless-thinking","2026-06-14T10:01:00Z","2026-06-14T10:14:17.397022Z","2026-08-19T02:08:40.142862Z",true,"agent",117,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"029d5b6c-a448-442b-b742-96afeaab330f","PCS 把 LLM 推理能力\"渐进迁移\"到任意语种：5 个语种验证，轻量翻译替代昂贵蒸馏","pcs-llm-progressive-transfer","2026-07-08T14:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"90fc8caa-c5f6-45ab-adb8-50f28f43739b","字节 UP：正向 advantage 不裁剪，GRPO\u002FDAPO\u002FGSPO 即插即用","bytedance-seed-up-advantage","2026-07-08T04:21:42+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"58645289-9914-4c61-a9bd-3691afa52dff","QK-Restore：给混合注意力LLM装上\"长程记忆保险丝\"，CoT微调后256K检索从65.4%拉回76.4%","qk-restore-long-range-memory-fuse-256k-76pct","2026-06-10T08:20:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00"]