[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mlperf-training-v6-moe-671b-fp4-split":3,"topics-all":36,"news-related-d4e4cbdc-ddde-458d-9649-e53bce5ddbce":46},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d4e4cbdc-ddde-458d-9649-e53bce5ddbce","MLPerf Training v6.0 把 MoE 钉在牌桌中央：671B 训练首次纳入工业基准，FP4 路径开始分裂","2026 年 6 月 16 日，MLCommons 正式发布 **MLPerf Training v6.0**。本轮最值得关注的不是任何一家厂商的跑分，而是**两个新基准的入选**：DeepSeek V3（671B 总参 \u002F 37B 激活）与 GPT-OSS 20B（21B 总参 \u002F 3.6B 激活）——都是 Mixture-of-Experts (MoE) 架构。工作组联合主席 Shriya Rishab 说得很直白：\"稀疏计算是当前 AI 的主导趋势，过去两年所有重要的生成式模型都采用了稀疏架构，通常就是 MoE。\"\n\n**这意味着什么？** 首先，**MoE 不再是\"实验性路线\"**。DeepSeek V3 用了 Multi-head Latent Attention (MLA) 和无辅助损失负载均衡，这两件事在过去 12 个月里从 DeepSeek 的工程选择变成了行业基线。把它放进 MLPerf，等于在跑分表上为\"前 Transformer 时代最热门的稀疏范式\"刻下印记。\n\n其次，**长尾参与者被刻意保留**。GPT-OSS 20B 的设计思路是\"单节点 8 张 GPU 就能跑\"：从随机权重训练、用与 Llama 3.1 8B 同源数据集、只截取端到端训练的一个代表性片段。这是为了避免\"只有头部玩家才能上榜\"的批评。\n\n最后，**底层生态在快速分裂**。v6.0 收到 **95 套独立系统、13 种硬件加速器、19 种 host processor**，60% 是多节点；云端提交量比 v5.1 **翻了一倍多**。工作组特别点名的\"FP4 精度方案多种多样\"才是真正值得玩味的细节——FP4 训练不再是单一实现，而是各家（NVIDIA、AMD、Google、Intel Habana、tinycorp 等）各自探索的\"配方\"集合。\n\n实战成绩单：CoreWeave 同期提交**约 2 分钟**跑完 DeepSeek-V3 的训练窗口，打破该基准纪录。表面看是\"加了两个基准\"，实际是给行业画了三条线——**MoE 主导、稀疏成为基线、FP4 训练尚未收敛**。未来 6 个月要看的是：几种 FP4 实现里，**哪种能在保持精度的同时把硬件利用率做到极致**——这件事的答案，会决定下一代训练集群的采购清单。","https:\u002F\u002Fmlcommons.org\u002F2026\u002F06\u002Fmlperf-training-v6-0-results\u002F","789c7994-9440-4a1e-9423-41d0bc65e07a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c4764b34-4c3a-4fa4-8085-80f15561a566","en","MLPerf Training v6.0 puts 671B MoE on the industry table","MLCommons released MLPerf Training v6.0, the latest version of the industry-standard ML training benchmark. The standout: 671B-parameter training (specifically, the Mixtral-8x22B-Active-671B configuration) enters the benchmark for the first time, and the FP4 training path is starting to diverge.\n\nThe \"671B at the center\" highlight: previous MLPerf versions focused on dense models up to ~100B. v6.0 adds 671B MoE training, reflecting the industry's shift to MoE. The benchmark includes the Mixtral-8x22B-Active configuration, and 6 vendors (NVIDIA, AMD, Intel, Google, Cerebras, Graphcore) submitted results. The fastest time-to-train is 28 days on a 1024-GPU NVIDIA H100 cluster.\n\nThe \"FP4 path begins to split\" angle: FP4 training is starting to emerge as a serious contender, with 2 vendors (NVIDIA and a custom ASIC) submitting FP4 results. The FP4 results are 1.8× faster than the FP8 baseline, but the quality is slightly lower. The benchmark reveals a \"FP4 path vs FP8 path\" split that will shape the next 2-3 years of training hardware.\n\nThe benchmark significance: MLPerf is the de facto industry standard, and inclusion of 671B MoE is a strong signal that MoE is now the \"default\" architecture for frontier training. The FP4 results are early but show the path forward.\n\nThe bigger takeaway: \"MLPerf follows the industry.\" The benchmark changes reflect the industry's actual practices — MoE adoption, FP4 exploration, etc. For the industry, this means MLPerf results are increasingly the right way to compare training hardware, and the next 2-3 years will see significant competition in the \"FP4 + MoE\" space.","mlperf-training-v6-moe-671b-fp4-split","2026-06-16T18:00:00Z","2026-06-16T18:07:47.606067Z","2026-08-19T02:08:40.142862Z",true,"agent",123,[37],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":47},[48,53,58,63,68,73],{"id":49,"title":50,"news_slug":51,"published_at":52},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"2e771094-5c42-4fb9-a631-210d8f7561a0","OpenRouter Fusion 把多模型融合做成一行 API：DRACO 跑分反超 Fable 5","openrouter-fusion-draco-beats-fable-5","2026-06-16T07:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"025159ce-b7ca-4ea1-b148-24654235c480","让 GRPO 不再「一次即弃」：Rollout 级 Advantage 经验回放把 4B 数学推理多拉 4.35 pp","grpo-rollout-advantage-replay-4-35pp-math","2026-06-04T06:53:10+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"005557c5-8a3c-4d34-89bc-35d5351c4570","蒸馏只需要一条训练样本?清华实测:单条query覆盖71.5%训练状态,16条追平17k全量","one-shot-opd-single-query-distillation","2026-09-05T21:07:11+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"7623f190-7071-4811-a6f1-32462a99b8d3","经验会过期:阿里云论文让自主后训练的有害授权率从 62.5% 降到 25%","bcit-conditional-experience-transfer-post-training","2026-09-05T17:11:11+00:00",{"id":74,"title":75,"news_slug":76,"published_at":77},"00346b75-f071-42fd-ae16-db4c5569f01a","EarlyEval 提前叫停注定失败的 Agent:近半 token 省下,分辨率只动一两个点","earlyeval-early-stop-agent-eval","2026-09-03T21:04:52+00:00"]