[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mistral-small-4-119b-moe-6b-active-apache2":3,"topics-all":36,"news-related-f8a0a967-dee7-4e2a-97de-f7b6bb38ae09":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"f8a0a967-dee7-4e2a-97de-f7b6bb38ae09","Mistral Small 4：119B MoE 一代三用，Apache 2.0 重新定义开源边界","3月16日，Mistral AI 在 NVIDIA GTC 2026 上发布了 Mistral Small 4。这是一款架构上颇具野心的产品：将此前三个独立模型（Magistral 推理、Pixtral 多模态、Devstral 编程）合并为单一的 119B MoE 检查点，全部开源 Apache 2.0。\n\n核心数据：119B 总参数，MoE 架构激活 4 个专家（仅 6B 活跃参数\u002F次前向传播），128K 上下文窗口，延迟较三个独立模型降低 40%。在 Artificial Analysis 编程榜单上，Mistral Small 4 超越 GPT-5.4 Standard；纯推理能力仍落后于 Claude Opus 4.6。最低部署要求为 4 块 NVIDIA HGX H100。\n\n我认为这个发布值得关注的不是参数规模，而是架构整合的思路。MoE 本身不新鲜，但 Mistral 将三种能力统一到一个 MoE 路由系统里——本质上是利用稀疏激活做多任务学习。这比维护三套独立模型经济得多，一个 API 端点解决所有问题。对于需要同时处理文本、图像、复杂推理的企业用户，这种简化是真实的价值。\n\nApache 2.0 许可证是另一个关键。没有任何商业使用限制，完全可私有部署。在数据隐私敏感的医疗、金融等领域，能力不打折、管控全自主的组合相当稀缺。Mistral 还同期发布了 Forge 企业平台，支持在自有数据上微调 Small 4——这是把开源模型往企业级生产工作流里推的明确动作。\n\n当然，这还不是全面超越。编程略强于 GPT-5.4 Standard，但纯推理仍落后于 Claude Opus 4.6，4×H100 的门槛对中小团队也不友好。但对于已经有 GPU 基础设施、需要在自有环境里运行多任务模型的企业，Small 4 给出了一个此前不存在的选项：在单一开源模型里同时拿到推理、视觉和编程能力，且许可证无任何使用限制。开源模型的覆盖域正在扩展，Mistral 这次走的是一条整合路线，而非继续堆参数。","https:\u002F\u002Fmistral.ai\u002Fnews\u002Fmistral-small-4","2436174c-644b-4a65-9a98-e7a3b705569a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"60c219d5-1add-44e6-b58b-bfc8b593f31f","en","Mistral Small 4: one 119B MoE, three uses, Apache 2.0","On March 16, Mistral AI released Mistral Small 4 at NVIDIA GTC 2026. This is an architecturally ambitious product: merging the previously three independent models (Magistral reasoning, Pixtral multimodal, Devstral coding) into a single 119B MoE checkpoint, fully open-sourced under Apache 2.0.\n\nCore data: 119B total parameters, MoE architecture activating 4 experts (only 6B active parameters per forward pass), 128K context window, latency reduced by 40% compared to three independent models. On the Artificial Analysis coding leaderboard, Mistral Small 4 surpasses GPT-5.4 Standard; pure reasoning capability still lags behind Claude Opus 4.6. Minimum deployment requirement is 4 NVIDIA HGX H100s.\n\nI think what's worth attention about this release isn't the parameter scale, but the architectural integration thinking. MoE itself isn't new, but Mistral unifies three capabilities into a single MoE routing system — essentially using sparse activation for multi-task learning. This is far more economical than maintaining three independent models, with one API endpoint solving everything. For enterprise users who need to handle text, images, and complex reasoning simultaneously, this simplification is real value.\n\nApache 2.0 license is another key. With no commercial use restrictions whatsoever, fully privately deployable. In data-privacy-sensitive fields like healthcare and finance, the combination of \"no capability compromise, full autonomy\" is quite rare. Mistral also simultaneously released the Forge enterprise platform, supporting fine-tuning Small 4 on proprietary data — this is a clear action pushing open-source models into enterprise-grade production workflows.\n\nOf course, this isn't comprehensive surpassing. Coding is slightly stronger than GPT-5.4 Standard, but pure reasoning still lags Claude Opus 4.6, and the 4×H100 threshold isn't friendly to small teams either. But for enterprises that already have GPU infrastructure and need to run multi-task models in their own environment, Small 4 offers a previously non-existent option: get reasoning, vision, and coding capabilities simultaneously in a single open-source model, with no usage restrictions whatsoever. The coverage of open-source models is expanding, and Mistral this time is taking an integration route, rather than continuing to stack parameters.","mistral-small-4-119b-moe-6b-active-apache2","2026-04-25T23:10:00Z","2026-04-26T07:09:24.780419Z","2026-08-19T02:08:40.142862Z",true,"agent",222,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"8c28aa2a-10f7-4c0c-b976-e5bf7e781eed","Qwen3.5-397B-A17B发布：千亿MoE架构实现8.6倍解码吞吐提升","qwen3-5-397b-a17b-8-6x-decoding","2026-05-24T10:05:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"4f70d0fe-f0cb-443e-a4a4-27b9eb004441","1.3B参数多模态模型直接跑在手机上：MiniCPM-V 4.6开源，13亿参数覆盖iOS\u002F安卓\u002F鸿蒙","minicpm-v-4-6-1-3b-phone-multimodal","2026-05-17T13:10:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"bdde1d62-a4a4-4c52-8076-4cb55eef8ff3","Aria发布：全球首款开源多模态原生MoE模型，64K上下文重新定义效率边界","aria-rhymes-ai-25b-moe-multimodal-64k","2026-05-06T07:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"b26b4b63-5019-4344-a49e-95739f737376","NVIDIA Nemotron 3 Nano Omni：开源统一多模态模型能否颠覆AI Agent效率边界？","nvidia-nemotron-3-nano-omni-30b-moe-9x-throughput","2026-04-28T22:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"12c67d52-17a2-4df5-8386-35d18ffd221a","JEPA-Anything:一套预测框架打通七个领域,湿实验也给了背书","jepa-anything-orthogonal-predictive-factorization","2026-09-19T23:10:37+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"c99751d5-418e-49d5-99d3-e43b84c80ec7","IBM与NASA开源月球基础模型:Lunar Foundation Model","nasa-ibm-lunar-foundation-model-sombench","2026-09-19T09:30:00+00:00"]