[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-over-thinking-test-time-compute-paradox":3,"news-related-103c151a-3ac1-42b0-92d2-d2f910dbc6ea":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"103c151a-3ac1-42b0-92d2-d2f910dbc6ea","大模型推理进入过思考时代：测试时计算的新问题","自OpenAI o1发布以来，测试时计算成了提升LLM能力的主流范式。但南京大学×百度等机构的最新研究揭示了一个关键悖论：**想太久反而会让模型答错。**\n\n研究首次系统性地挑战了推理越长效果越好的假设。通过边际收益曲线分析，研究者发现随推理token增加收益递减显著。更关键的是过思考（Overthinking）现象——模型在加长推理链时会意外抛弃之前正确的中间答案，最终给出错误结论。\n\n研究者还发现：**最优思考长度与题目难度高度相关**。简单问题在较低预算就达到负边际收益，而难题需要更长的推理链。这意味着均匀分配推理预算是一种次优策略。\n\n研究提出的成本感知评估框架显示，在中等推理预算处停止推理，可大幅降低计算量同时保持相近准确率。换句话说，**少想一点，不仅省钱，效果可能还更好**。\n\n当行业还在卷模型参数量时，一个更精细的问题已浮现：LLM需要学会知道什么时候该停止思考。这也将推动自适应推理预算分配、动态停止机制等工程优化方向。","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2604.10739v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"31a78eaa-ce1f-4c39-b76a-28bc4e3c6896","en","The over-thinking era: new questions in test-time compute","arXiv 2604.10739v1 investigates the \"over-thinking\" phenomenon in test-time compute: when given a simple problem, reasoning models still generate thousands of tokens of \"thinking\" before answering, with the extra compute not improving and sometimes hurting accuracy. The paper proposes methods to detect and skip unnecessary thinking steps.","over-thinking-test-time-compute-paradox","2026-06-01T19:00:00Z","2026-06-01T19:06:00.946244Z","2026-08-19T02:08:40.142862Z",true,"agent",129,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7d371b09-9792-465d-b73a-3d0af4735129","InferenceBench：15 个前沿 Agent 自主做 LLM 推理优化","inferencebench-open-ended-llm-optimization","2026-08-16T12:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"c71b8ee7-9487-4c78-89fd-30bb0368b99e","DeepSeek V4 Flash：284B\u002F13B MoE，成本比 Luna 低 60%","deepseek-v4-flash-0731-intelligence-index-50","2026-08-05T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"c48681ff-ddbb-402c-ade8-23b584a06aea","更强教师反而教不动学生：Lightning OPD 2.0 剥掉蒸馏中的“文风噪声”","lightning-opd-2-cross-teacher-style-bias","2026-07-30T16:17:15+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"beec1ff3-22af-4657-b58a-90cb0797c3b1","PyroDash 让小模型「借力」大模型推理：把 LLM 调用砍到 1.9%，成本从 $49 降到 $1.78","pyrodash-small-large-routing","2026-07-24T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"7f723663-8405-43aa-b31e-73efd714fa97","KV-Cache Grafting：冻结权重，Gemma-4-12B AIME 80%→93.3%","byte-exact-kv-cache-grafting","2026-07-17T06:20:00+00:00"]