[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-xiaomi-mimo-v2-6-open-weights":3,"topics-all":38,"news-related-0bc17892-3c47-4081-86a4-3d90afa0c54b":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"0bc17892-3c47-4081-86a4-3d90afa0c54b","小米 MiMo-V2.6 开源:万亿 MoE 追平 Grok 4.7,Flash 三分之一价格保九成战力","小米开源 MiMo-V2.6 系列:旗舰 Pro 为 1.02T 总参\u002F42B 激活稀疏 MoE,1M 上下文原生全模态,MIT 协议。Artificial Analysis 智能指数 46 分追平 Grok 4.7,创开源权重最高;Flash 310B\u002F15B 以约三分之一价格保留九成基准战力。","9 月 22 日,小米 MiMo 团队发布并开源 MiMo-V2.6 系列。旗舰 MiMo-V2.6-Pro-RL 以 1.02 万亿总参数、420 亿激活参数的稀疏 MoE 架构,把开源权重模型在 Artificial Analysis 智能指数上推到 46 分——与 xAI 前一天发布的 Grok 4.7 并列,是该榜单口径下开源权重模型的最高分。\n\n## 架构:万亿稀疏 MoE + 1M 上下文 + 原生全模态\n\n按官方模型卡,MiMo-V2.6-Pro-RL 总参数 1.02T,每个 token 激活 42B,路由专家 384 个、每 token 激活 8 个;骨干 70 层,其中 60 层滑窗注意力(SWA)、10 层全局注意力,hidden size 6144。上下文长度 1M token,原生支持文本、图像、视频、音频四种输入:视觉侧是 681M 参数的 MiMo ViT,音频侧是 308M AudioTokenizer 加 127M 音频 patch encoder,另配 5 层 MTP 投机解码器加速推理。权重以 MIT 协议在 Hugging Face 开放,允许商用(官方模型卡见 [Hugging Face](https:\u002F\u002Fhuggingface.co\u002FXiaomiMiMo\u002FMiMo-V2.6-Pro-RL))。\n\n## 评测:Agent 基准贴身闭源旗舰,安全项跨代跳变\n\n模型卡的对比表里,MiMo-V2.6-Pro-RL 在多个 Agent 基准上与闭源旗舰互有胜负:DeepSWE v1.1 拿 71.9,超过 Claude Fable 5 的 70.0,略低于 Claude Opus 5 的 74.0 与 GPT-5.6 Sol 的 73.0;AutomationBench 53.1,高于 Opus 5(50.3)、Sol(45.8)和 Fable 5(46.2);Agents' Last Exam 31.6 与 Opus 5 持平;Terminal Bench 2.1 拿 89.9;JobBench 62.0,明显高于 Sol 的 45.4。安全方向进步最猛:CyberGym 从上一代 V2.5 Pro 的 40.0 跳到 94.0,自家的 MiMo Cyber Bench 从 0.0 直接拉到 80.2。\n\n作为对照,V2.5 Pro 在 DeepSWE v1.1 只有 19.0、AutomationBench 只有 16.0——V2.6 在 Agent 与安全两条线上都是跨数量级的代际跳变。\n\n第三方口径上,Artificial Analysis 给 MiMo-V2.6-Pro 打出智能指数 46 分,为其榜上开源权重模型最高分,与 Grok 4.7 并列(见 [Artificial Analysis](https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fmimo-v2-6-pro))。小米官方口径是 46.32 分,并称其为\"最强开源模型\"、超过 Kimi K3 与 Qwen3.8 Max——官方宣传,读者自行加权。\n\n## 训练:You Only RL Once + 组内智能体评分\n\nV2.6 的主线是把强化学习往自我改进推。官方称训练采用\"一次混合 RL\"(You Only RL Once):coding、通用 Agent、视觉、网络安全任务混进同一个 RL run,不同 harness 上的能力互相迁移。底层是全异步 GRPO,每步 1,568 个 prompt、每个 prompt 16 条 rollout。\n\n更有意思的是打分方式:二元的 pass\u002Ffail 无法给\"都通过\"的解排序,于是 MiMo 引入组内智能体评分(Groupwise Agentic Grading)——离线用组内对比 rollout 生成任务特定 rubric(Groupwise Reward Synthesis),在线给通过的轨迹排序并把 advantage 导向更优解(Groupwise Advantage Redistribution),相当于把奖励信号本身也做了 scaling。配套手段还包括冷启动自纠错、环境加固与 verifier 交叉检查防 reward hacking,RL 之后再做多前缀多教师 on-policy 蒸馏(MOPD2)。\n\n## Flash:310B 版本才是行业影响所在\n\n同系列 MiMo-V2.6-Flash-RL 总参数 310B、激活 15B,基准却掉得不多:DeepSWE 67.9 对 Pro 的 71.9,AutomationBench 52.3 对 53.1,Terminal Bench 2.1 87.6 对 89.9,普遍保住九成以上。[VentureBeat](https:\u002F\u002Fventurebeat.com\u002Ftechnology\u002Fbetter-than-deepseek-xiaomis-mimo-v2-6-pro-debuts-as-the-top-open-weights-model-in-the-world-alongside-cheaper-v2-6-flash) 报道中,bitsandbytes 作者、CMU 教授 Tim Dettmers 的评价是:Flash\"感觉是 300B 到 550B 级别里最好的模型,比 DeepSeek v4.1 和 GLM 5.3 Flash 都好\"。API 定价上,Pro 为每百万 token 输入 0.435 美元、输出 0.87 美元,Flash 只要 0.14 与 0.28 美元——大约三分之一的价格。\n\n当万亿旗舰的能力能以三分之一价格在 15B 激活的模型上保住九成,开源阵营对闭源 API 的价格压力就会从\"旗舰对决\"下沉到\"日常调用\"。登顶是头条,Flash 才是账单。","https:\u002F\u002Fhuggingface.co\u002FXiaomiMiMo\u002FMiMo-V2.6-Pro-RL","fcab25c6-b3c7-4150-947f-def2ae1ef01e",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"3eea7264-8321-4556-b764-0aae3dca8121","en","Xiaomi MiMo-V2.6: 1T MoE ties Grok 4.7, Flash keeps 90% at 1\u002F3 price","Xiaomi open-sources MiMo-V2.6: 1.02T\u002F42B MoE Pro ties Grok 4.7 at 46 on the Artificial Analysis Index; 310B Flash keeps 90% of benchmarks at 1\u002F3 price.","On September 22, the Xiaomi MiMo team released and open-sourced the MiMo-V2.6 series. The flagship MiMo-V2.6-Pro-RL — a sparse mixture-of-experts model with 1.02 trillion total parameters and 42 billion activated per token — scores 46 on Artificial Analysis's Intelligence Index, tying xAI's Grok 4.7 released the day before, and marking the highest open-weights result on that leaderboard.\n\n## Architecture: Trillion-Parameter Sparse MoE, 1M Context, Native Omnimodal\n\nPer the official model card, MiMo-V2.6-Pro-RL has 1.02T total parameters with 42B active per token: 384 routed experts, 8 activated per token, a 70-layer backbone (60 sliding-window attention layers plus 10 global-attention layers), and a hidden size of 6144. Context length is 1M tokens, with native support for text, image, video, and audio input: a 681M-parameter MiMo ViT on the vision side, a 308M AudioTokenizer plus a 127M audio patch encoder on the audio side, and a 5-layer MTP speculative decoder for faster inference. Weights are available on Hugging Face under the MIT license, which permits commercial use (official model card: [Hugging Face](https:\u002F\u002Fhuggingface.co\u002FXiaomiMiMo\u002FMiMo-V2.6-Pro-RL)).\n\n## Benchmarks: Neck-and-Neck with Closed Frontiers on Agent Eval, a Generational Jump on Security\n\nIn the model card's comparison table, MiMo-V2.6-Pro-RL trades blows with closed-source flagships on agent benchmarks: 71.9 on DeepSWE v1.1, above Claude Fable 5's 70.0 and just below Claude Opus 5's 74.0 and GPT-5.6 Sol's 73.0; 53.1 on AutomationBench, ahead of Opus 5 (50.3), Sol (45.8), and Fable 5 (46.2); 31.6 on Agents' Last Exam, level with Opus 5; 89.9 on Terminal Bench 2.1; and 62.0 on JobBench, clearly above Sol's 45.4. The security track shows the sharpest jump: CyberGym leaped from V2.5 Pro's 40.0 to 94.0, and Xiaomi's own MiMo Cyber Bench went from 0.0 to 80.2.\n\nFor context, V2.5 Pro scored just 19.0 on DeepSWE v1.1 and 16.0 on AutomationBench — V2.6 is a generational leap of an order of magnitude on both the agent and security lines.\n\nOn third-party scoring, Artificial Analysis gives MiMo-V2.6-Pro a 46 on its Intelligence Index, the top open-weights result on its leaderboard, tied with Grok 4.7 (see [Artificial Analysis](https:\u002F\u002Fartificialanalysis.ai\u002Fmodels\u002Fmimo-v2-6-pro)). Xiaomi's own claim is 46.32 and \"the strongest open-source model to date,\" surpassing Kimi K3 and Qwen3.8 Max — that's vendor framing, weight it accordingly.\n\n## Training: You Only RL Once, plus Groupwise Agentic Grading\n\nThe through-line of V2.6 is scaling reinforcement learning toward self-improvement. Xiaomi says training used \"You Only RL Once\": one mixed RL run spanning coding, general agents, visual, and cybersecurity tasks, so capabilities reinforce each other and transfer across harnesses. The substrate is fully asynchronous GRPO — 1,568 prompts with 16 rollouts each per step.\n\nThe more interesting part is the grading. Binary pass\u002Ffail cannot rank passing solutions, so MiMo introduces Groupwise Agentic Grading: offline, Groupwise Reward Synthesis (GRS) builds task-specific rubrics from contrasting rollouts within each group; online, Groupwise Advantage Redistribution (GAR) ranks passing trajectories and shifts advantage toward higher-quality solutions — effectively scaling the reward signal itself. The loop is kept honest with a self-correction cold start, environment hardening, adversarial screening, and verifier cross-checks against reward hacking, followed by multi-prefix multi-teacher on-policy distillation (MOPD2) after RL.\n\n## Flash: The 310B Sibling Is Where the Industry Impact Lands\n\nThe sibling MiMo-V2.6-Flash-RL has 310B total and 15B active parameters, yet holds most of the ground: 67.9 vs Pro's 71.9 on DeepSWE, 52.3 vs 53.1 on AutomationBench, 87.6 vs 89.9 on Terminal Bench 2.1 — generally retaining more than 90% of Pro's scores. In [VentureBeat](https:\u002F\u002Fventurebeat.com\u002Ftechnology\u002Fbetter-than-deepseek-xiaomis-mimo-v2-6-pro-debuts-as-the-top-open-weights-model-in-the-world-alongside-cheaper-v2-6-flash)'s coverage, Tim Dettmers, creator of bitsandbytes and a CMU professor, called Flash \"the best model in the 300B to 550B class. Better than DeepSeek v4.1 and GLM 5.3 Flash.\" On API pricing, Pro runs $0.435 per million input tokens and $0.87 per million output tokens, while Flash costs $0.14 and $0.28 — roughly a third of the price.\n\nWhen a trillion-parameter flagship's capability can be retained at 90%-plus by a 15B-active model at a third of the price, open-weight competition against closed APIs shifts from \"flagship showdowns\" down to \"everyday calls.\" The top spot is the headline; Flash is the bill.","xiaomi-mimo-v2-6-open-weights","2026-09-22T13:02:37Z","2026-09-22T13:11:46.548029Z","2026-09-22T13:11:46.548038Z",true,"agent",9,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"21fe3c11-4ba4-4801-b6fc-60c4ae559dc1","Yandex 逆流开源:35B 参数的 T5 MoE,每个 token 只激活 0.6B","yandex-aliceai-t5-sparse-moe","2026-09-16T19:11:43+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"8e730a3d-439b-45cf-961d-f77cf01469fd","Cohere 开源 218B 翻译专用 MoE:25B 激活,自测评分超 DeepL,2×H100 可部署","cohere-north-small-translate","2026-09-11T19:07:20+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"33f3b08b-c8a2-43ec-81cf-85e2b918f913","腾讯开源 Hy4 preview:770B MoE、1M 上下文,模型首次参与自身训练","tencent-hy4-preview-770b-moe","2026-08-29T15:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"3d36921f-3b84-4663-97a0-fee7d4eff795","汤森路透开源 Thomson-1.0-Small:持续学习改造 Qwen,3B 激活的 35B MoE","thomson-1-0-small-continual-learning","2026-08-28T19:10:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00"]