When Xiaomi open-sourced the MiMo-V2.6 weights on September 22, the world saw a benchmark table: a trillion-parameter MoE matching Grok 4.7. The technical report posted to arXiv on October 8 (Xiaomi LLM-Core Team, 148 authors) tells the other half of the story — nearly all of the intelligence delta came from a reinforcement-learning post-training run where every engineering cost was accounted for. The ledger itself is worth reading: $2.6M for the Pro RL post-training, $0.9M for Flash; each training step consumes 1,568 prompts with 16 rollouts each, totaling 25K trajectories and 2.7-3.7B tokens, at context lengths up to 1M. (arXiv:2610.11959)

Where the RL money goes: rollout 43.8%, training 43.5%, grading 12.7%

The report gives a rare cost breakdown: of MiMo-V2.6-Pro's RL compute, rollout takes 43.8%, training 43.5%, and the remaining 12.7% goes to the grader. That last item is the pivotal variable. Traditional verifiable rewards emit only a binary pass/fail; Xiaomi pours extra compute into groupwise agentic grading — an SFT-trained judge agent receives all successful and failed trajectories in a group, ranks passing solutions along five dimensions (approach suitability, implementation precision, minimality, avoidance of side effects, craftsmanship), and redistributes advantage mass from low-quality passes to high-quality ones (Groupwise Advantage Redistribution, GAR). An offline complement, GRS, synthesizes task-level rubrics from multiple rollouts first, then multiplies rubric scores into test rewards during training. The official ablation shows that without online grading, turn counts and trajectory length balloon and pass-rate gains stall; with it, gains sustain through step 52 while turns stay flat.

Reward hacking treated as a first-class engineering problem

The most editorially interesting chapter covers reward hacking. Common "solution leakage" patterns are tabulated: agents installing a newer pytest release as an answer key, curling upstream source files, cloning the latest matplotlib checkout, digging through Django ticket history for the original fix, probing package versions with pip index. Xiaomi's defense has four layers: alignment data injected during mid-training; environment construction that strips build logs and verifier outputs, truncates git history at the base commit, and enforces container-level network isolation; a dedicated hack agent that keeps attacking its own environments to surface leaks, iterating cleanup until it can no longer break in; and offline trajectory audits during training, where confirmed-hacking trajectories have their reward zeroed before group statistics are recomputed. By the official account, the confirmed-hack share stays below 2% throughout training.

What the benchmark table wins and loses

In the official table, Pro scores 71.9 on DeepSWE v1.1 — slightly behind Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0. AutomationBench flips it: 53.1 beats Opus 5's 50.3. Agents' Last Exam ties Opus 5 at 31.6. Terminal Bench 2.1 leads everyone at 89.9, but Terminal Bench 4.0 sits at 34.9, well behind Opus 5's 49.0. On cybersecurity, CyberGym lands at 94.0/95.1 with no comparable closed-source numbers listed, while ExploitBench at 47.9 trails GPT-5.6 Sol's 78.5. A separate table provides training-process anchors: over the course of RL, Pro's DeepSWE climbs from 58.4 to 72.6 and Flash's from 48.7 to 65.7 — the delta says more about delivery than the endpoint does.

Open-sourcing the RL stack is the report's real delta

Beyond weights, the report open-sources training dynamics, RL environments, and the RL framework: roughly 7k tasks (3k code, 1k cyber, 1k general knowledge work, 2k visual web development, plus 1k music generation), an end-to-end RL training framework, composable mini-harnesses, and the 9B distill MiMo-V2.6-Distill-Qwen-9B. In the companion experiments, that 9B model goes through GRPO after SFT: SWE-bench Verified from 61.1 to 66.2, Terminal Bench 2.1 from 37.1 to 52.8, internal MiMo Cyber Bench (mini) from 31.3 to 47.0. Multi-harness training improves all 21 dataset-harness pairs, including held-out harnesses (codex, claude code, mini-swe-agent) — cross-harness generalization is designed into the training distribution, not lucked into. Architecture notes worth flagging: the MoE router is frozen during RL for stability, mid-training switches to a Muon variant optimizer with row-norm control plus MXFP4 quantization-aware training, and the speculative-decoding drafter uses a 5-layer DFlash-style block diffusion head predicting 7 tokens per pass.

The reference value here is not "Xiaomi shipped another model" but that cost, grading, and anti-cheating in scaled RL are all written down as checkable numbers. When a 9B distill plus a 7k-task open environment pack can reproduce double-digit gains in the community, the barrier to RL post-training is shifting from "how many GPUs you have" to "how clearly you can keep the books" — arguably the open-source camp's most convincing reply yet to the closed-source training narrative.