Sampling temperature usually lives on the inference side of an LLM stack: dial it up for more diverse outputs, down for more conservative ones. A paper submitted to arXiv on September 27 drags that knob into the training loop. TGRL (Temperature-Grouped Reinforcement Learning), from a Meituan and Chinese Academy of Sciences team, turns the contrast between low- and high-temperature rollouts into an explicit training signal. The authors report that at the same rollout budget, RLVR training reaches equivalent accuracy up to 36% faster, with CodeForces rating up 196.7 points. The paper has been accepted as a NeurIPS 2026 poster, and the code is open-sourced.
Exploration is the expensive part of RLVR
Reinforcement learning with verifiable rewards (RLVR) is the dominant recipe for training reasoning models: math and code problems come with checkable answers, so reward signals are clean. Exploration efficiency, however, remains a central bottleneck. To expose a student model to more diverse solutions, the standard moves are raising the sampling temperature or test-time scaling. The paper points out that both paths either expand the rollout budget at sampling time — cost scales up — or leave the benefit of exploration unquantified. Higher temperature does produce more diverse trajectories, but the optimizer has no way to know which of that diversity deserves reward, or which tokens should get the credit.
The method: group by temperature, then credit by token
TGRL works in two layers. First, for each prompt, the rollout group is partitioned into a low-temperature subset and a high-temperature subset; the reward contrast between the two estimates this round's exploration gain — the extra reward the high-temperature group earns over the low-temperature one is direct evidence that exploration paid off. Second, that group-level signal is allocated as token-level credit: for the same logits, the method computes the Jensen-Shannon divergence between the two temperature-scaled next-token distributions. Positions where divergence is largest are exactly where temperature genuinely changed the decision, so that is where the credit lands. Crucially, the pipeline never expands the rollout budget — the grouping reuses rollouts that were going to be sampled anyway.
The numbers: 11 benchmarks, 36% faster training
The authors report results across 11 benchmarks: at 32B scale, the six-benchmark math average improves by 1.6%, CodeForces rating rises by 196.7 points, LiveCodeBench Pass@16 by 4.4%, and ALFWorld/WebShop success rates by 6.3% and 4.9%. The mechanism ablation shows a clean two-stage gain: on Qwen3-14B, a high-temperature GRPO baseline scores 67.0; adding mixed-temperature grouping brings 68.2; adding JS credit assignment (full TGRL) reaches 69.4. On Qwen3-32B the ladder is 66.2 → 67.8 → 70.2. Note that all of these numbers are self-reported by the authors in the paper and repository; no independent third-party replication exists yet.
Signals on the engineering side
The codebase is built on the open-source verl distributed RL framework; the repository's dependency list includes SGLang-oriented and NPU-oriented files, and the contact addresses span Meituan and the Institute of Automation, CAS. The author list spans the University of Chinese Academy of Sciences, Meituan, and MAIS & NLPR at CAS's Institute of Automation, with two equal-contribution first authors and two corresponding authors — seven authors in total. For a method accepted as a NeurIPS 2026 poster, open code plus mainstream-framework integration means the replication barrier is low; third-party replications are what to watch next.
For teams already running RLVR pipelines, the takeaway is direct: sampling temperature does not have to be a mere hyperparameter. It can be treated as a free source of diversity, quantified and fed back into training. While the industry pays real compute for exploration, squeezing a second helping of value out of samples you already paid for may be the cheapest efficiency story available.
Paper: arXiv:2609.33589 (https://arxiv.org/abs/2609.33589). Code: github.com/1229095296/TGRL