[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-xiaomi-mimo-ultraspeed-latency-pricing":3,"topics-all":38,"news-related-5409a0b8-8d54-4b20-9310-8cde3239b132":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"5409a0b8-8d54-4b20-9310-8cde3239b132","小米把推理速度卖成商品:同一权重,十倍价","小米 MiMo-V2.6-Pro-UltraSpeed 用与标准版相同的 1T 参数权重,通过推理侧加速把输出速度拉高约十倍,价格也定在十倍。本文拆解「速度层」定价:同 checkpoint、同 1M 上下文,变的只是 batching 与投机解码等服务端工程,对批处理任务不值,对交互式 Agent 循环却可能值。","9 月 21 日,小米把 MiMo-V2.6 系列的强化学习权重传上 Hugging Face,第二天 UltraSpeed 服务层上线,从无日期、无定价、无 API 到全部就位只用了约 48 小时。这一周里真正值得注意的不是又一家厂商发了万亿参数模型,而是它把「推理速度」单独拿出来标了价:同样的权重,十倍的价格,换来大约十倍的输出速度。\n\n## 同一权重,两种卖法\n\nUltraSpeed 不是小模型、不是蒸馏版,也不是重新训练的版本。小米官方描述它就是旗舰 MiMo-V2.6-Pro 的快速版,基于同一个 1T checkpoint 构建,质量与原版一致。第三方目录与 OpenRouter 的页面也是同样的口径:同 checkpoint、同 1M token 上下文窗口、同文本\u002F图像\u002F视频\u002F音频原生多模态输入。发生变化的是服务侧——batching、投机解码、硬件配额、并发限制,这些介于权重和你的 HTTP 请求之间的工程。一个服务优化变体的存在,说明小米相信有一批买家愿意为 token 延迟付比 token 本身更多的钱;它不说明模型变强了,也不说明权重变了。\n\n## 十倍速度的两种说法\n\n小米官方材料声称 UltraSpeed 达到标准 Pro 服务的「最高 20 倍」输出速度;而 OpenRouter 与多家第三方目录在同一天的描述是「约 10 倍」。两个数字同时在流通,没有人公布过能调和它们的测量。诚实的读法是:「最高 20 倍」是厂商在未指明条件下的上限,大概率是有利 batch size 下的最好情况;「约 10 倍」是目录愿意断言的典型值。在有人公布特定并发水平下的逐请求延迟分布之前,真实倍数应视为一个较宽区间。方向没有争议:这个层级确实明显更快,而且定价也确实按「快」来收钱。\n\n## 算一笔交互循环的账\n\n价格比与速度比几乎对齐:输入 $4.35\u002F百万 token、输出 $8.70\u002F百万 token,对比标准 Pro 的 $0.44\u002F$0.87,输入约 9.9 倍、输出整整 10 倍。这笔账怎么算,完全取决于延迟附着在什么上。批处理任务——离线文档处理、评测、数据集生成——延迟几乎一文不值,标准 Pro 通道是显然的选择。但交互式 Agent 循环里每一轮都卡着下一轮,延迟会复利:30 步的 Agent 运行中,10 倍速的模型端到端不是 10 倍速,但它是「用户等得下去」和「用户放弃」的区别。OpenRouter 的实测数据可以佐证需求真实性:UltraSpeed P50 吞吐 107 tok\u002Fs,调用方前五名是 Kilo Code、Hermes Agent、omp、Cursor、pi——清一色编码 Agent 与 Agent 框架,单 Kilo Code 一家就送来 20.2B token。\n\n## 权重照旧 MIT,服务另算账\n\n小米对 MiMo-V2.6 RL 权重挂的是 MIT 许可证,技术报告与部署说明随包发布,还开源了 RL 训练环境与代码。这比一次 API 发布披露得更多。但 MIT 标签只覆盖已发布的权重:UltraSpeed 是托管服务,受服务条款而非权重许可约束。在自己硬件上跑这份 checkpoint 的权利,和转售别人对它的加速托管的权利,是两种不同的权利,只有前者带着 MIT 标签。对想压成本的人,答案一直没变:权重可以下载,标准 Pro 通道 $0.44\u002F$0.87,是调用 1T 级模型最便宜的方式之一。\n\n值得盯的是三件事:独立并发延迟测量会告诉我们真实倍数靠近 10 还是 20;UltraSpeed 的第一次价格调整会告诉我们这是永久溢价还是首发定价——以 10 倍起价的层级很少停在 10 倍。对读者来说,问题其实很简单:如果你的产品就是墙钟时间,买它;如果你期待的是更好的模型,别买。速度层卖的是时间,不是智能。","https:\u002F\u002Fwww.orcarouter.ai\u002Fblog\u002Fxiaomi-mimo-v2-6-pro-ultraspeed","f88344b2-8136-4f83-9c49-310b753c2bd3",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"821b38ba-d564-4334-a196-d0f0300968b9","en","Xiaomi Sells Inference Speed: Same Weights, 10x Price","Xiaomi's MiMo-V2.6-Pro-UltraSpeed serves the same 1T checkpoint ~10x faster at 10x the price — worthless for batch jobs, decisive for agent loops.","On September 21, Xiaomi uploaded the MiMo-V2.6 reinforcement-learning weights to Hugging Face; the UltraSpeed serving tier went live the next day. Going from no date, no pricing and no API to fully shipped took roughly 48 hours. The noteworthy part of this release is not another trillion-parameter model — it is that Xiaomi put a price tag on inference speed alone: identical weights, ten times the price, for roughly ten times the output speed.\n\n## Same weights, two price lists\n\nUltraSpeed is not a smaller model, a distilled model, or a differently trained one. Xiaomi describes it as the fast edition of the flagship MiMo-V2.6-Pro, built from the same 1T checkpoint, matching the original in quality. Third-party catalogues and the OpenRouter page carry the same story: same checkpoint, same 1M-token context window, same native multimodal inputs across text, image, video and audio. What changes is the serving side — batching, speculative decoding, hardware allocation, concurrency limits: the engineering between the weights and your HTTP request. The existence of a serving-optimised variant tells you Xiaomi believes a class of buyers values token latency more than token cost. It does not tell you the model got better, and it does not tell you the weights changed.\n\n## Two versions of the 10x claim\n\nXiaomi's own material claims UltraSpeed reaches up to 20x the output speed of the standard Pro service; OpenRouter and several third-party catalogues, describing the same model on the same day, say roughly 10x. Both numbers are in circulation, and nobody has published a measurement that reconciles them. The honest reading: \"up to 20x\" is an unspecified vendor ceiling, probably at favourable batch sizes, while \"roughly 10x\" is what a catalogue was willing to assert as typical. Until someone publishes per-request latency distributions at a stated concurrency, treat the real multiplier as a wide band. The direction is not in dispute: the tier is materially faster, and it is priced as if it were.\n\n## The arithmetic of an interactive loop\n\nPrice ratio and speed ratio nearly align: $4.35 per million input tokens and $8.70 per million output, against $0.44\u002F$0.87 for standard Pro — about 9.9x on input, exactly 10x on output. Whether that trade is worth taking depends on what the latency is attached to. Batch jobs — offline document processing, evaluation runs, dataset generation — get nearly zero value from latency; the standard Pro lane is the obvious call. But in an interactive agent loop, every turn gates the next, and latency compounds: across a 30-step agent run, a 10x-faster model is not 10x faster end to end, but it is the difference between a session a user waits through and one they abandon. OpenRouter's live data supports the demand: UltraSpeed's P50 throughput is 107 tok\u002Fs, and its top callers are Kilo Code, Hermes Agent, omp, Cursor and pi — coding agents and agent frameworks across the board, with Kilo Code alone sending 20.2B tokens.\n\n## Weights stay MIT, the service does not\n\nXiaomi tagged the MiMo-V2.6 RL weights MIT, shipped a technical report and deployment notes, and open-sourced the RL training environment and code. That is broader disclosure than an API launch. But the MIT tag covers only the published weights: UltraSpeed is a hosted service governed by terms of service, not by a weight licence. The right to run the checkpoint on your own hardware and the right to resell someone's accelerated serving of it are different rights — only the first comes with the MIT tag. For anyone optimizing cost, the answer has not changed: the weights are downloadable, and the standard Pro lane at $0.44\u002F$0.87 remains one of the cheapest ways to call a 1T-class model.\n\nThree things are worth watching: an independent latency measurement at stated concurrency would tell you whether the real multiplier sits nearer 10x or 20x; the first price move on UltraSpeed would tell you whether this is a permanent premium or a launch number — a tier that starts at 10x rarely stays there. For readers the question is simple: if wall-clock time is your product, buy it; if you expected a better model, don't. A speed tier sells time, not intelligence.","xiaomi-mimo-ultraspeed-latency-pricing","2026-10-09T15:11:05Z","2026-10-09T15:13:03.897619Z","2026-10-09T15:13:03.897626Z",true,"agent",19,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"44b118b2-cb3d-47bb-8f26-bbce777cfb31","1 亿 rollout 背后:Beam 的 RL 训练工厂","beam-rl-factory-100m-rollouts","2026-10-07T07:15:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"f695ab3f-b8cf-4d0a-9707-71deed44069c","Qwen和GLM把旗舰权重卖出双倍价","qwen-glm-prime-speed-tier","2026-09-26T21:10:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"2bd44b6f-5688-471f-930e-17a93984e8e7","中国电信开源星辰 Xing4.0:昇腾全栈训练的 29B MoE","xing4-29b-a4b-ascend-moe","2026-09-19T15:10:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"0fe869ca-11c4-4831-afb1-19a31fd88dfc","智谱公开国内大模型首个 RSI:GLM-5.3 Infra Agent 在 10 万国产卡集群自建推理,2 周吞吐 3 倍","zhipu-glm-rsi-infrastructure-chinese-cluster","2026-09-17T08:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"1a50eda4-e62d-40ba-8f9d-dab756067e2d","16GB 内存跑 313B GLM-5.3-Flash:WARP 把专家权重搬进 NVMe","warp-engine-glm-flash-nvme-inference","2026-08-31T13:00:00+00:00"]