What does RL post-training actually buy you? A paper posted to arXiv on October 1 (arXiv:2610.01509, carrying a Meta badge on its Hugging Face paper page) offers an awkward answer: it may merely be sharpening behaviors the base model already had — better single-shot accuracy, at the cost of shrinking solution coverage. The authors name this hidden cost the "Sharpening Tax," and provide both a way to measure it and a fix.
The counter-intuitive finding: base models with a light harness can fight
The starting point is a popular hypothesis: RL post-training only hones existing base-model behaviors, raising pass@1 while depressing pass@K. This trade-off had been repeatedly observed on math and coding tasks, and the default assumption was that it would carry over to agentic tasks — multi-turn tool use looks like it should depend on capabilities newly acquired during post-training.
The measurements say otherwise. The authors find that a pre-trained LLM equipped with a light inference harness can serve as a capable agent: despite far lower pass@1, given enough test-time budget the base model's pass@K often surpasses its post-trained counterpart. In other words, the expensive RL post-training you paid for may be a net liability on the "eventually solvable under repeated sampling" axis.
The mechanism: tasks pushed to two extremes
Why does this happen? The paper's analysis shows post-training pushes tasks toward two extremes — always solved or never solved — buying sampling efficiency and consistency at the price of solution coverage. One side casts a wide net and eventually catches fish; the other tightens the mesh, more accurate per cast but leaking more of the catch.
To quantify the loss, the authors propose the Sharpening Tax metric, measuring how much test-time scalability is lost after post-training. The experimental scale is substantial: 14 base/post-trained model pairs from four families, three agentic benchmarks, 42 cases in total. The tax turns out to be prevalent in most settings, estimable from just a few rollouts, and well correlated with other metrics.
PTGS: temperature per prompt, less tax paid
Diagnosis without a prescription would be incomplete. The paper closes with posterior-tempered group sampling (PTGS), a plug-and-play Bayesian sampler that adapts the sampling temperature to each prompt's estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline — solving more tasks under repeated sampling while also improving single-shot accuracy.
So what
The paper stings three audiences differently. Benchmark folks: pass@1 alone systematically overstates the value of post-training; agentic evaluations should add a pass@K view. Practitioners choosing models: if your workload tolerates repeated sampling (batch jobs, offline tasks, retry loops with verifiers), a base model plus a light harness is a baseline worth running seriously — it may not lose to the official post-trained release. Post-training teams: the sampling temperature in your RL recipe is not optional noise; it directly determines how much tax you pay — per-difficulty tempering schemes like PTGS belong in the default configuration.
Paper: https://arxiv.org/abs/2610.01509 (HF page: https://huggingface.co/papers/2610.01509)