[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-sharpening-tax-rl-post-training":3,"topics-all":38,"news-related-4c4e444f-9614-42fe-b8f4-743204f854fa":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"4c4e444f-9614-42fe-b8f4-743204f854fa","RL 后训练收「锐化税」:base 模型配轻 harness,pass@K 反超官方版","Meta 团队提出「锐化税」:RL 后训练把任务推向「总能解出或永远解不出」两端,单发更准、解覆盖缩水。base 模型配轻 harness 后 pass@1 虽低,给足采样预算 pass@K 常反超后训练版。14 组模型对、3 个 agentic 基准共 42 案例里这笔税普遍存在;PTGS 按题调温可少交。","RL 后训练到底给模型带来了什么？Meta 团队 10 月 1 日挂在 arXiv 上的新论文（arXiv:2610.01509）给了一个不太体面的答案：它可能只是在「锐化」基座模型已有的行为——单次采样更准，解的覆盖面却在缩水。论文把这笔隐性代价命名为「锐化税」（Sharpening Tax），并给出了度量方法与补救方案。\n\n## 反直觉的发现：base 模型配个轻 harness 就能打\n\n论文的出发点是一个流行假说：RL 后训练只是把基座模型已有的行为磨得更锋利，提升 pass@1 的代价是 pass@K 下降。这个权衡此前在数学和代码任务上被反复观察到，大家默认它也会延续到 agentic 任务——毕竟多轮工具调用看起来更依赖后训练新学的能力。\n\n实测结果正好相反。作者发现，预训练 LLM 配上一个轻量推理 harness，就能当个像样的 agent：虽然 pass@1 低不少，但只要给够测试时预算，base 模型的 pass@K 经常反超对应的后训练版本。换句话说，你花大价钱做的 RL 后训练，在「反复采样总能解出」这个维度上可能是负资产。\n\n## 税的机制：任务被推向两个极端\n\n为什么会这样？论文的机制分析给出解释：后训练把任务推向「总能解出」和「永远解不出」两个极端，以此换来采样效率和一致性，代价就是解的覆盖面。一边是广撒网总能捞到几条，一边是网眼变密、单网更准但漏掉的鱼更多。\n\n为了量化这笔损失，作者提出 Sharpening Tax 指标，度量后训练之后测试时可扩展性的损失。实验规模不小：14 组 base\u002Fpost-trained 模型对、4 个模型家族、3 个 agentic 基准，共 42 个案例。结论是这笔税在大多数设置下普遍存在，而且从少量 rollout 就能估出来，与其他指标的相关性也不错。\n\n## PTGS：按题调温，少交点税\n\n光诊断不开药不算完整。论文最后给出 posterior-tempered group sampling（PTGS）：一个即插即用的贝叶斯采样器，按每条 prompt 的估计难度自适应调整采样温度。在两个 agentic 环境的 RL 训练里，PTGS 比固定温度基线交的税更少——重复采样能解出更多任务，单发准确率也一并提升。\n\n## 所以呢\n\n这篇论文对三类人各有刺痛感。做评测的：只看 pass@1 会系统性高估后训练的价值，agentic 场景该补上 pass@K 视角。做选型的：如果你的场景能承受多次采样（批处理、离线任务、带验证器的重试回路），base 模型加轻 harness 是个值得认真跑的基线，未必输给官方后训练版。做后训练的：RL 配方里的采样温度不是可有可无的超参，它直接决定你交多少税——PTGS 这类按难度调温的策略，值得进默认配置。\n\n论文地址：https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.01509（HF 论文页：https:\u002F\u002Fhuggingface.co\u002Fpapers\u002F2610.01509）","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.01509","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"726b0016-b4dd-4a3a-8ac8-2da6ea925695","en","Sharpening Tax: RL post-training shrinks pass@K coverage","An arXiv paper shows RL post-training polarizes tasks and shrinks pass@K coverage: base models plus a light harness often beat post-trained twins.","What does RL post-training actually buy you? A paper posted to arXiv on October 1 (arXiv:2610.01509, carrying a Meta badge on its Hugging Face paper page) offers an awkward answer: it may merely be sharpening behaviors the base model already had — better single-shot accuracy, at the cost of shrinking solution coverage. The authors name this hidden cost the \"Sharpening Tax,\" and provide both a way to measure it and a fix.\n\n## The counter-intuitive finding: base models with a light harness can fight\n\nThe starting point is a popular hypothesis: RL post-training only hones existing base-model behaviors, raising pass@1 while depressing pass@K. This trade-off had been repeatedly observed on math and coding tasks, and the default assumption was that it would carry over to agentic tasks — multi-turn tool use looks like it should depend on capabilities newly acquired during post-training.\n\nThe measurements say otherwise. The authors find that a pre-trained LLM equipped with a light inference harness can serve as a capable agent: despite far lower pass@1, given enough test-time budget the base model's pass@K often surpasses its post-trained counterpart. In other words, the expensive RL post-training you paid for may be a net liability on the \"eventually solvable under repeated sampling\" axis.\n\n## The mechanism: tasks pushed to two extremes\n\nWhy does this happen? The paper's analysis shows post-training pushes tasks toward two extremes — always solved or never solved — buying sampling efficiency and consistency at the price of solution coverage. One side casts a wide net and eventually catches fish; the other tightens the mesh, more accurate per cast but leaking more of the catch.\n\nTo quantify the loss, the authors propose the Sharpening Tax metric, measuring how much test-time scalability is lost after post-training. The experimental scale is substantial: 14 base\u002Fpost-trained model pairs from four families, three agentic benchmarks, 42 cases in total. The tax turns out to be prevalent in most settings, estimable from just a few rollouts, and well correlated with other metrics.\n\n## PTGS: temperature per prompt, less tax paid\n\nDiagnosis without a prescription would be incomplete. The paper closes with posterior-tempered group sampling (PTGS), a plug-and-play Bayesian sampler that adapts the sampling temperature to each prompt's estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline — solving more tasks under repeated sampling while also improving single-shot accuracy.\n\n## So what\n\nThe paper stings three audiences differently. Benchmark folks: pass@1 alone systematically overstates the value of post-training; agentic evaluations should add a pass@K view. Practitioners choosing models: if your workload tolerates repeated sampling (batch jobs, offline tasks, retry loops with verifiers), a base model plus a light harness is a baseline worth running seriously — it may not lose to the official post-trained release. Post-training teams: the sampling temperature in your RL recipe is not optional noise; it directly determines how much tax you pay — per-difficulty tempering schemes like PTGS belong in the default configuration.\n\nPaper: https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.01509 (HF page: https:\u002F\u002Fhuggingface.co\u002Fpapers\u002F2610.01509)","sharpening-tax-rl-post-training","2026-10-03T15:09:45Z","2026-10-03T15:09:53.223226Z","2026-10-03T15:09:53.223235Z",true,"agent",930,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"c368ad9f-9308-4a0b-8f5c-3ae4601b48b9","D2K-Bench: 专家设计把 LLM 写 GPU 核提速 33.9%","d2k-bench-llm-gpu-kernel-design-guidance","2026-10-07T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"85f7f1c2-d896-436b-a915-37faed8776eb","你改主意了,模型没改:被拒需求也会带偏大模型","intent-eval-rejected-change-confusion","2026-10-06T17:15:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"206ea36a-1eea-463f-a241-1e3b32f5ec2d","USTC GraphForge:证据图把任务和 rubric 钉在一起,Qwen3.6-27B 涨三基准","graphforge-ustc-qwen36-27b-evidence-graph","2026-10-04T03:05:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"c8684f9e-1028-4b81-a530-b04b0e044d3b","DiscoBench：搜得越多反而越错？首个聚焦\"何时该向用户问清楚\"的搜索 Agent 基准","discobench-clarify-search","2026-07-05T10:15:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"75908d93-a928-4695-b966-d98847d135cb","SkillComposer 把 Agent 的技能选择重做成技能序列生成:GPT-5.2-Codex 提升 +23.1pp","skillcomposer-skill-sequence","2026-07-01T04:15:00+00:00"]