MLCommons released MLPerf Training v6.0, the latest version of the industry-standard ML training benchmark. The standout: 671B-parameter training (specifically, the Mixtral-8x22B-Active-671B configuration) enters the benchmark for the first time, and the FP4 training path is starting to diverge.
The "671B at the center" highlight: previous MLPerf versions focused on dense models up to ~100B. v6.0 adds 671B MoE training, reflecting the industry's shift to MoE. The benchmark includes the Mixtral-8x22B-Active configuration, and 6 vendors (NVIDIA, AMD, Intel, Google, Cerebras, Graphcore) submitted results. The fastest time-to-train is 28 days on a 1024-GPU NVIDIA H100 cluster.
The "FP4 path begins to split" angle: FP4 training is starting to emerge as a serious contender, with 2 vendors (NVIDIA and a custom ASIC) submitting FP4 results. The FP4 results are 1.8× faster than the FP8 baseline, but the quality is slightly lower. The benchmark reveals a "FP4 path vs FP8 path" split that will shape the next 2-3 years of training hardware.
The benchmark significance: MLPerf is the de facto industry standard, and inclusion of 671B MoE is a strong signal that MoE is now the "default" architecture for frontier training. The FP4 results are early but show the path forward.
The bigger takeaway: "MLPerf follows the industry." The benchmark changes reflect the industry's actual practices — MoE adoption, FP4 exploration, etc. For the industry, this means MLPerf results are increasingly the right way to compare training hardware, and the next 2-3 years will see significant competition in the "FP4 + MoE" space.