[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tgrl-temperature-grouped-rl":3,"topics-all":38,"news-related-9dc3bd1e-95da-4c07-b352-900688d1af44":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"9dc3bd1e-95da-4c07-b352-900688d1af44","TGRL把温度分组变成训练信号:RLVR训练提速36%","美团与中科院团队提出温度分组强化学习 TGRL:把同一 prompt 的 rollout 按高低温度分组,用两组奖励差估计探索增益,再经 JS 散度落到 token 级信用。论文称同等采样预算下训练最多快 36%,CodeForces 提升 196.7 分,已被 NeurIPS 2026 接收,代码开源。","温度这个参数,在大多数 LLM 部署里只是个推理侧旋钮:调高一点,输出更发散;调低一点,更保守。9 月 27 日提交到 arXiv 的一篇论文把它拽进了训练回路——美团与中国科学院团队提出的 TGRL(温度分组强化学习)用「高低温度对比」直接产出训练信号,论文称同等采样预算下,RLVR 训练达到同等精度最多快 36%,CodeForces 评分提升 196.7 分。论文已被 NeurIPS 2026 接收为 Poster,代码同步开源。\n\n## 探索,RLVR 最贵的那部分开销\n\n可验证奖励强化学习(RLVR)是当前推理模型训练的主流配方:数学题、代码题有标准答案,奖励信号清晰。但探索效率一直是瓶颈。想让学生模型见到更多样的解法,常规手段是调高采样温度或做 test-time scaling——论文指出,这两条路要么在采样时直接扩大 rollout 预算(训练成本随之上涨),要么让「探索到底有没有用」停留在不可量化的状态。温度调高了,轨迹确实更多样,但优化器并不知道这些多样性里哪些值得奖励、该把功劳记在哪个 token 头上。\n\n## 方法:先按温度分组,再落到 token 信用\n\nTGRL 的做法分两层。第一层,同一个 prompt 的 rollout 组被划分成低温度子集和高温度子集,两组的奖励对比用来估计这一轮的探索增益——高温组比低温组多赚到的奖励,就是「探索值得做」的直接证据。第二层,这个组级信号被分配成 token 级信用:对同一份 logits 按两种温度缩放出 next-token 分布,计算两者的 Jensen-Shannon 散度,散度大的位置正是温度真正改变决策的地方,信用就落在那里。整条链路没有扩大 rollout 预算——分组复用的本来就是原本要采样的那些 rollout。\n\n## 数字:11 个基准,同预算快 36%\n\n作者报告的结果覆盖 11 个基准:32B 规模上六个数学基准平均分提升 1.6%,CodeForces 评分提升 196.7 分,LiveCodeBench Pass@16 提升 4.4%,ALFWorld\u002FWebShop 成功率分别提升 6.3% 和 4.9%。机制消融呈现清晰的两段式增益:Qwen3-14B 上,高温 GRPO 基线 67.0 分,加入混合温度分组后 68.2,再加 JS 信用分配(TGRL 完整版)到 69.4;Qwen3-32B 上是 66.2 → 67.8 → 70.2。需要说明,以上数字均为作者在论文与代码仓库中的自报口径,目前还没有第三方独立复现。\n\n## 工程侧的信号\n\n代码基于开源分布式 RL 框架 verl 构建,仓库依赖清单里同时有 SGLang 相关与 NPU 相关的文件,联系方式给出美团邮箱与中科院自动化所邮箱。作者阵容横跨中国科学院大学、美团与中科院自动化所 MAIS & NLPR,两位共同一作、两位通讯作者,论文署名共 7 人。对一个被 NeurIPS 2026 收为 Poster 的方法来说,开源加上主流框架适配意味着复现门槛不高,接下来值得盯的就是第三方复现结果。\n\n对正在跑 RLVR 训练的团队,这篇论文的启示很直接:采样温度不必只是一个超参,它可以被当成免费的多样性来源,量化之后反哺训练。当整个行业都在为探索付出真金白银的算力时,把「已经花掉的采样」榨出第二份价值,可能是当下最划算的效率故事。\n\n原文:arXiv:2609.33589(https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.33589),代码:github.com\u002F1229095296\u002FTGRL","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.33589","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"04291b31-ef09-4795-89e6-9cf2e3bb181a","en","TGRL Turns Temperature Into a Training Signal: RLVR 36% Faster","TGRL turns temperature contrast into token-level credit: same rollout budget, 36% faster RLVR, +196.7 CodeForces, NeurIPS 2026.","Sampling temperature usually lives on the inference side of an LLM stack: dial it up for more diverse outputs, down for more conservative ones. A paper submitted to arXiv on September 27 drags that knob into the training loop. TGRL (Temperature-Grouped Reinforcement Learning), from a Meituan and Chinese Academy of Sciences team, turns the contrast between low- and high-temperature rollouts into an explicit training signal. The authors report that at the same rollout budget, RLVR training reaches equivalent accuracy up to 36% faster, with CodeForces rating up 196.7 points. The paper has been accepted as a NeurIPS 2026 poster, and the code is open-sourced.\n\n## Exploration is the expensive part of RLVR\n\nReinforcement learning with verifiable rewards (RLVR) is the dominant recipe for training reasoning models: math and code problems come with checkable answers, so reward signals are clean. Exploration efficiency, however, remains a central bottleneck. To expose a student model to more diverse solutions, the standard moves are raising the sampling temperature or test-time scaling. The paper points out that both paths either expand the rollout budget at sampling time — cost scales up — or leave the benefit of exploration unquantified. Higher temperature does produce more diverse trajectories, but the optimizer has no way to know which of that diversity deserves reward, or which tokens should get the credit.\n\n## The method: group by temperature, then credit by token\n\nTGRL works in two layers. First, for each prompt, the rollout group is partitioned into a low-temperature subset and a high-temperature subset; the reward contrast between the two estimates this round's exploration gain — the extra reward the high-temperature group earns over the low-temperature one is direct evidence that exploration paid off. Second, that group-level signal is allocated as token-level credit: for the same logits, the method computes the Jensen-Shannon divergence between the two temperature-scaled next-token distributions. Positions where divergence is largest are exactly where temperature genuinely changed the decision, so that is where the credit lands. Crucially, the pipeline never expands the rollout budget — the grouping reuses rollouts that were going to be sampled anyway.\n\n## The numbers: 11 benchmarks, 36% faster training\n\nThe authors report results across 11 benchmarks: at 32B scale, the six-benchmark math average improves by 1.6%, CodeForces rating rises by 196.7 points, LiveCodeBench Pass@16 by 4.4%, and ALFWorld\u002FWebShop success rates by 6.3% and 4.9%. The mechanism ablation shows a clean two-stage gain: on Qwen3-14B, a high-temperature GRPO baseline scores 67.0; adding mixed-temperature grouping brings 68.2; adding JS credit assignment (full TGRL) reaches 69.4. On Qwen3-32B the ladder is 66.2 → 67.8 → 70.2. Note that all of these numbers are self-reported by the authors in the paper and repository; no independent third-party replication exists yet.\n\n## Signals on the engineering side\n\nThe codebase is built on the open-source verl distributed RL framework; the repository's dependency list includes SGLang-oriented and NPU-oriented files, and the contact addresses span Meituan and the Institute of Automation, CAS. The author list spans the University of Chinese Academy of Sciences, Meituan, and MAIS & NLPR at CAS's Institute of Automation, with two equal-contribution first authors and two corresponding authors — seven authors in total. For a method accepted as a NeurIPS 2026 poster, open code plus mainstream-framework integration means the replication barrier is low; third-party replications are what to watch next.\n\nFor teams already running RLVR pipelines, the takeaway is direct: sampling temperature does not have to be a mere hyperparameter. It can be treated as a free source of diversity, quantified and fed back into training. While the industry pays real compute for exploration, squeezing a second helping of value out of samples you already paid for may be the cheapest efficiency story available.\n\nPaper: arXiv:2609.33589 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.33589). Code: github.com\u002F1229095296\u002FTGRL","tgrl-temperature-grouped-rl","2026-10-01T19:02:31Z","2026-10-01T19:08:22.455210Z","2026-10-01T19:08:22.455217Z",true,"agent",93,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"2e27016d-b90e-45c7-825a-41fd1e435c80","JHU 新研究:组合持续学习机制,百任务记忆留存从 1.2% 提到 34.9%","compose-cl-long-horizon-memorization","2026-09-16T15:10:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"2731ed1c-17c3-4d85-9174-983cf50743e3","地铁售票机上的 AI 大考:2.6GB 端侧模型 91.32 分超 GPT-5.6,规则基线也拿 84.6","metrollm-bench-transit-kiosk-llm","2026-09-12T23:08:18+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"9ebb888c-dfe7-416a-9940-a913527d4f73","AI Agent 的失败比成功更值钱:5 万对错误诊断数据,修正通过率 18.4%→51.1%","agent-error-dataset","2026-10-01T15:11:08+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"3e28cfaf-74aa-4521-8eba-37332fe93901","7块ESP32跑1.58-bit Qwen:功耗1.53瓦","esp32s3-bitnet-llm-cluster","2026-09-30T13:11:01+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"ccce6dfe-776f-4cd8-9605-9163daea4627","Jeff 开源决策模型:2B 追平 Jev","jeff-open-decision-models-0-8b","2026-09-29T15:10:00+00:00"]