On September 21, Xiaomi uploaded the MiMo-V2.6 reinforcement-learning weights to Hugging Face; the UltraSpeed serving tier went live the next day. Going from no date, no pricing and no API to fully shipped took roughly 48 hours. The noteworthy part of this release is not another trillion-parameter model — it is that Xiaomi put a price tag on inference speed alone: identical weights, ten times the price, for roughly ten times the output speed.
Same weights, two price lists
UltraSpeed is not a smaller model, a distilled model, or a differently trained one. Xiaomi describes it as the fast edition of the flagship MiMo-V2.6-Pro, built from the same 1T checkpoint, matching the original in quality. Third-party catalogues and the OpenRouter page carry the same story: same checkpoint, same 1M-token context window, same native multimodal inputs across text, image, video and audio. What changes is the serving side — batching, speculative decoding, hardware allocation, concurrency limits: the engineering between the weights and your HTTP request. The existence of a serving-optimised variant tells you Xiaomi believes a class of buyers values token latency more than token cost. It does not tell you the model got better, and it does not tell you the weights changed.
Two versions of the 10x claim
Xiaomi's own material claims UltraSpeed reaches up to 20x the output speed of the standard Pro service; OpenRouter and several third-party catalogues, describing the same model on the same day, say roughly 10x. Both numbers are in circulation, and nobody has published a measurement that reconciles them. The honest reading: "up to 20x" is an unspecified vendor ceiling, probably at favourable batch sizes, while "roughly 10x" is what a catalogue was willing to assert as typical. Until someone publishes per-request latency distributions at a stated concurrency, treat the real multiplier as a wide band. The direction is not in dispute: the tier is materially faster, and it is priced as if it were.
The arithmetic of an interactive loop
Price ratio and speed ratio nearly align: $4.35 per million input tokens and $8.70 per million output, against $0.44/$0.87 for standard Pro — about 9.9x on input, exactly 10x on output. Whether that trade is worth taking depends on what the latency is attached to. Batch jobs — offline document processing, evaluation runs, dataset generation — get nearly zero value from latency; the standard Pro lane is the obvious call. But in an interactive agent loop, every turn gates the next, and latency compounds: across a 30-step agent run, a 10x-faster model is not 10x faster end to end, but it is the difference between a session a user waits through and one they abandon. OpenRouter's live data supports the demand: UltraSpeed's P50 throughput is 107 tok/s, and its top callers are Kilo Code, Hermes Agent, omp, Cursor and pi — coding agents and agent frameworks across the board, with Kilo Code alone sending 20.2B tokens.
Weights stay MIT, the service does not
Xiaomi tagged the MiMo-V2.6 RL weights MIT, shipped a technical report and deployment notes, and open-sourced the RL training environment and code. That is broader disclosure than an API launch. But the MIT tag covers only the published weights: UltraSpeed is a hosted service governed by terms of service, not by a weight licence. The right to run the checkpoint on your own hardware and the right to resell someone's accelerated serving of it are different rights — only the first comes with the MIT tag. For anyone optimizing cost, the answer has not changed: the weights are downloadable, and the standard Pro lane at $0.44/$0.87 remains one of the cheapest ways to call a 1T-class model.
Three things are worth watching: an independent latency measurement at stated concurrency would tell you whether the real multiplier sits nearer 10x or 20x; the first price move on UltraSpeed would tell you whether this is a permanent premium or a launch number — a tier that starts at 10x rarely stays there. For readers the question is simple: if wall-clock time is your product, buy it; if you expected a better model, don't. A speed tier sells time, not intelligence.