[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-nemotron-imo-gold-open-recipe":3,"topics-all":38,"news-related-54b86d93-0fd0-4107-9353-9b79a1446f69":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"54b86d93-0fd0-4107-9353-9b79a1446f69","NVIDIA 开源 IMO 金牌完整配方:30\u002F42 分、561B 双专家、算力账本全公开","NVIDIA 用 Nemotron 3 Ultra 三个 checkpoint 组成的纯自然语言管线,在 IMO 2026 官方拿到 30\u002F42 分,越过 29 分金线。论文放出两个 561B 后训练专家模型、训练数据、推理代码和 200 题新基准,连 23.1 亿 token、4800 GPU 小时的算力账也公开。","IMO 2026 赛场上,NVIDIA 的 Nemotron 系统作为正式参赛者拿到 30\u002F42 分,超过 29 分的金牌线,全部答卷由官方阅卷人评分:P1、P2、P4、P5 拿满分,P3 和 P6 各得 1 分。与依赖形式化证明器的 AlphaProof 路线不同,这套系统从头到尾只用自然语言——不用形式化证明器、不调外部工具、不联网,吃进官方 LaTeX 题面,吐出人类可读的证明。9 月 9 日团队把整条管线写成技术报告挂上 arXiv([arXiv:2609.10712](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10712)),金牌配方可复现。\n\n## 三个 checkpoint,一套 16 票全票制\n\n管线骨架是迭代式「生成—验证—精炼」。生成侧由三个 Nemotron 3 Ultra checkpoint 组成:公开基座 GA,加上分别经 SFT 和 RL 后训练的两个专家。第一轮每题生成 384 个证明尝试——每个 checkpoint 从 8 个互补提示词各采样 16 次,提示词分别引导引理分解、路线对比、反例搜索等不同解题策略。\n\n验证侧由 RL、SFT 两个 checkpoint 充当免参考验证器:每个证明收 16 个独立判分,满分 1 分、小错 0.5 分、致命错误 0 分,16 票全部给满才算「接受」。无人过线时,系统从证明池挑排名前 16 的候选,附上最多 8 条验证批评送回三个模型精炼,每轮 192 次新尝试,最多迭代 8 轮。\n\n## 后训练的增益有硬数字\n\n30 题开发集上,两个后训练专家全面压过基座:8 轮搜索加兜底后,RL 管线累计 180 分、SFT 165 分,GA 基座只有 162 分(独立评审团口径)。预算怎么花更划算,论文给出了直接对照:RL 单模型把尝试加到 256 次只接受 14 题;换成 RL 与 SFT 各 128 次的混合池,接受题数涨到 18。团队对此的结论是:第二个 checkpoint 能解出第一个解不出的题,而给单个 checkpoint 翻倍采样增益甚微——换模型比堆采样值钱。\n\n## 算力账本:7 亿 token 找齐全部证明\n\n四个满分证明在开赛后 76 分钟内全部通过终选面板,另两题也在 100 分钟内定稿;全部 6 个提交证明约耗 7.07 亿 token、1464 个 GB200 GPU 小时,把在途轮次跑满则总计约 23.1 亿 token、4800 GPU 小时。赛后复盘留了个警示:切断时刻两套模型验证器都给出约 32 分的估计,比官方评分高 2 分,差额全部来自 P3 和 P6——模型评审在同一批证明上集体看走眼,这是模型验证的共性盲区,不是某一家的噪声。\n\n## 五件套全开源,561B 权重直接下载\n\nHF collection([nvidia\u002Fnemotron-labs-imo-2026](https:\u002F\u002Fhuggingface.co\u002Fcollections\u002Fnvidia\u002Fnemotron-labs-imo-2026))里躺着两个 561B 专家 checkpoint(OpenMDW-1.1 许可)、13.3 万条 SFT 语料加 9600 条 RL 题集(CC BY 4.0),以及与 Titu Andreescu 教授合作的 200 题全新基准 Nemotron-IMO-Bench;推理代码进 NeMo-Skills,RL 配方进 NeMo-RL,连提交的证明原文都在。论文结论值得抄下来:光堆生成规模不够,收益来自互补的后训练 checkpoint、验证制导精炼,以及花在终选上的算力——更花哨的路由和激进过滤,反而会把后来被证明正确的候选提前扔掉。\n\n对做推理系统的人,这份配方的价值不在分数本身,而在每一步都带消融数字和算力账,可以直接对答案。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10712","474eef8c-e0c3-46cf-adee-c089558220f9",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1addd41b-6468-4205-b5cf-6a971990b35f","en","NVIDIA Open-Sources Its IMO Gold Recipe: 561B Ensemble Weights","NVIDIA's Nemotron pipeline scored 30\u002F42 at IMO 2026, clearing the gold line with no formal prover. Weights, data, code and a new benchmark are open.","At IMO 2026, NVIDIA's Nemotron system competed officially and scored 30 out of 42, one point above the 29-point gold cutoff, with every submission graded by official IMO graders: full credit on Problems 1, 2, 4 and 5, one point each on Problems 3 and 6. Unlike the AlphaProof line of work that leans on formal provers, this system runs entirely in natural language — no formal prover, no external tools, no internet. It takes the official LaTeX problem statements and produces human-readable proofs. On September 9 the team posted the full pipeline as a technical report ([arXiv:2609.10712](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10712)), turning the gold recipe into something reproducible.\n\n## Three checkpoints and a 16-vote unanimity rule\n\nThe backbone is an iterative generate-verify-refine loop. Generation uses three Nemotron 3 Ultra checkpoints: the general-availability base model plus two specialists post-trained with SFT and RL respectively. Round one produces 384 proof attempts per problem — each checkpoint samples 16 attempts from each of eight complementary prompts that steer different strategies such as lemma-first decomposition, route comparison, or counterexample search.\n\nVerification is handled by the RL and SFT checkpoints acting as reference-free verifiers. Every proof receives 16 independent judgments, scoring 1 for complete and correct, 0.5 for minor errors, 0 for fatal ones; a proof is accepted only when all 16 judgments award a full score. If nothing passes, the system picks the top-16 proofs from the pool, attaches up to eight verifier critiques, and sends them back for refinement — 192 fresh attempts per round, at most eight rounds.\n\n## The post-training gain, in hard numbers\n\nOn the 30-problem development set both specialists beat the base model: after eight rounds plus fallback, the RL pipeline reaches 180 cumulative points and SFT 165, against 162 for GA (independent-jury scoring). The paper also weighs how to spend the search budget: a single RL model pushed to 256 attempts accepts 14 problems, while a mixed pool of RL 128 + SFT 128 accepts 18. The team's takeaway: a second checkpoint solves problems the first cannot, and doubling attempts of a single checkpoint adds little.\n\n## The compute ledger\n\nThe four full-credit proofs cleared final selection within 76 minutes of the contest start, the remaining two within 100 minutes. All six submitted proofs were found within roughly 707M generated tokens and 1,464 GB200 GPU-hours; completing the in-flight rounds brought the full run to about 2.31B tokens and 4,800 GPU-hours. The post-hoc analysis carries a warning: at the cutoff both model-based verifiers estimated roughly 32 points, two above the official 30, with the gap entirely on Problems 3 and 6 — model juries shared a blind spot rather than noising independently.\n\n## Everything open, 561B weights downloadable\n\nThe HF collection ([nvidia\u002Fnemotron-labs-imo-2026](https:\u002F\u002Fhuggingface.co\u002Fcollections\u002Fnvidia\u002Fnemotron-labs-imo-2026)) holds the two 561B specialist checkpoints (OpenMDW-1.1 license), 133k SFT samples plus 9.6k RL problems (CC BY 4.0), and Nemotron-IMO-Bench, 200 novel olympiad problems built with Professor Titu Andreescu; inference code lives in NeMo-Skills, the RL recipe in NeMo-RL, submitted proofs included. The conclusion is worth quoting: scaling proof generation alone is not enough — the gains came from complementary post-trained checkpoints, verification-guided refinement that preserves candidates, and compute spent on final selection, while fancier routing and aggressive filtering discarded candidates that later turned out correct.\n\nFor anyone building reasoning systems, the value is not the 30 points themselves but that every step ships with ablation numbers and a compute bill you can check against.","nvidia-nemotron-imo-gold-open-recipe","2026-09-11T17:13:27Z","2026-09-11T17:13:30.589419Z","2026-09-11T17:13:30.589427Z",true,"agent",151,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"2731ed1c-17c3-4d85-9174-983cf50743e3","地铁售票机上的 AI 大考:2.6GB 端侧模型 91.32 分超 GPT-5.6,规则基线也拿 84.6","metrollm-bench-transit-kiosk-llm","2026-09-12T23:08:18+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"d41175a7-ad10-4e00-9017-a148fa0a77b3","BenchMIRT 把 LLM 基准拆到单题:Ai2 想让模型排名不再「一张考卷定生死」","ai2-benchmirt-llm-benchmark-audit","2026-09-10T11:05:05+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"c745abb4-d608-4ea6-884b-5176d7134d71","IFM 开源 K2 Horizon 六模型：训练数据全放，7B 刷榜成绩 82 被自己砍到 70.6","ifm-k2-horizon-open-fleet-audit","2026-09-05T23:07:55+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"42b7939c-1b44-43b8-95cf-a8fc2204560d","NVIDIA 开源 Personal AI Router，把家里 RTX 与 Mac 拼成本地 AI 集群","nvidia-personal-ai-router-pair-beta","2026-09-04T03:20:00+00:00"]