[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mit-rllm-idle-cycle-2x-train-half-energy":3,"topics-all":36,"news-related-3a9a8c69-d668-4d2c-ae82-caeba45aa2d5":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"3a9a8c69-d668-4d2c-ae82-caeba45aa2d5","MIT新方法利用计算空闲周期：推理模型训练速度翻倍，能耗减半","推理大语言模型（RLLM）通过逐步分解复杂问题来得出答案，在高级编程和多步规划等任务上表现出色。然而，其训练过程却面临严重的效率瓶颈：MIT 研究团队发现，在强化学习训练中，生成多个候选答案的 rollout 阶段占据了高达 85% 的执行时间，而模型权重更新这一真正的训练部分反而耗时甚少。当部分高性能处理器忙于生成候选答案时，其他处理器只能处于空闲等待状态，造成算力的巨大浪费。\\n\\n针对这一问题，MIT 与 NVIDIA、ETH Zurich、MIT-IBM Watson AI Lab 及 UMass Amherst 的联合团队提出了一种自适应训练方法：用一个更小更快的辅助模型来预测主推理模型的输出，再由主模型验证这些预测。当某些处理器空闲时，辅助模型接管其算力；当主模型需要验证时，辅助模型暂停工作。这种自适应调度机制确保了 GPU 集群中的每一块芯片都不会被闲置，在不损失精度的情况下将训练速度提升了一倍，同时降低了能耗和成本。\\n\\n这一成果的更大意义在于，它揭示了当前 RL 训练范式的一个系统性缺陷——当行业普遍追求更大参数、更多算力时，训练流程本身的效率问题往往被忽视。对行业而言，这一突破的启示是：更高效的 RL 训练方法意味着未来可以用更少的资源训练出更强能力的推理模型；训练系统本身的优化可能是下一阶段 AI 进步的关键杠杆。","https:\u002F\u002Fnews.mit.edu\u002F2026\u002Fnew-method-could-increase-llm-training-efficiency-0226","4613a0c2-8d14-4485-b855-f8fad33c4527",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"201afca5-c77f-47aa-a186-45131cac63ea","en","MIT taps idle compute cycles: 2x training, half the energy","Reasoning large language models (RLLMs) work by progressively decomposing complex problems to reach answers, performing well on advanced coding and multi-step planning tasks. But their training process faces a serious efficiency bottleneck: an MIT research team found that in RL training, the rollout phase generating multiple candidate answers occupies up to 85% of execution time, while the actual model weight update — the real training part — takes relatively little. When some high-performance processors are busy generating candidate answers, others can only sit idle, wasting huge amounts of compute.\n\nTo address this, a joint team from MIT, NVIDIA, ETH Zurich, the MIT-IBM Watson AI Lab, and UMass Amherst proposed an adaptive training method: a smaller, faster auxiliary model predicts the main reasoning model's output, and the main model verifies those predictions. When some processors are idle, the auxiliary model takes over their compute; when the main model needs to verify, the auxiliary model pauses. This adaptive scheduling ensures no chip in the GPU cluster sits idle, doubling training speed without losing accuracy, while reducing energy consumption and cost.\n\nThe broader significance of this work: it exposes a systemic flaw in the current RL training paradigm. As the industry universally pursues larger parameters and more compute, the efficiency of the training process itself is often ignored. For the industry, the takeaway is this: more efficient RL training methods mean future reasoning models can be trained with fewer resources; optimization of the training system itself may be the key lever for the next stage of AI progress.","mit-rllm-idle-cycle-2x-train-half-energy","2026-05-22T08:10:00Z","2026-05-22T16:09:05.366271Z","2026-08-19T02:08:40.142862Z",true,"agent",213,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"217f417d-1b9c-475b-99f4-e21e7c909711","MHAR 把 Transformer 残差流从「单车道」拆成 H 条独立路由:子空间第一次有权自己挑历史层","multi-head-attention-residuals-mhar","2026-08-01T07:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"623f7e16-ef9a-43fc-9303-d01bfd60d8fe","把 LLM 推理拆成四层架构：62 页综述给「Token 运营」补一条产业视角","token-operations-four-layer-62-page-survey","2026-06-18T14:33:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"4e43e35d-a808-4125-be31-69cadedc61f1","PoLar 把 LLM 层变成可调积木：动态跳层+复读，3B 模型数学推理涨 60+ 个百分点","polar-icml-2026-3b-math-62pp-jump","2026-06-15T14:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"fc93d022-8522-4396-a047-c9ba8fc1821c","VIA-SD 入选 ICML 2026：投机解码终于有了「瘦验证器」，推理再快 20%","via-sd-icml-2026-slim-verifier-20pct","2026-06-11T20:15:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"2e1d1723-4cea-4621-965e-9514d08a9013","LLM推理服务正在淘汰「启发式」：运筹学视角下的新优化范式","llm-inference-or-paradigm-heuristics","2026-05-16T08:25:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"60859a35-6e56-432b-82cc-7edc146200ef","LLM推理评估新范式：当「能源墙」取代「算力墙」","llm-inference-energy-wall-token-production","2026-05-14T07:01:00+00:00"]