[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ai2-emo-emergent-modular-moe-pluggable":3,"news-related-461816c1-61b7-4077-b482-428379c046de":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"461816c1-61b7-4077-b482-428379c046de","AI2 EMO：把 MoE 训练成「可拆装」模块，1B 激活也能按域调度","现在大多数 MoE 模型都号称「激活少量参数」，但实际部署时仍要把整张卡加载进去。原因并不神秘：标准 MoE 的专家在训练中往往只学会区分「介词」「专有名词」这种浅层词法特征，而不是「医学」「代码」这种语义域。一旦想只保留部分专家，模型性能就会断崖式下跌。\\n\\nAI2 这篇工作换了一个更聪明的角度：把「涌现模块化」做成预训练的一阶目标。方法看似简单——同一文档的所有 token 共享一个由路由器挑出的专家子集——但配合「全局负载均衡」与训练时随机采样子集大小这两个工程细节后，效果立竿见影：1B 激活的 14B MoE 在保留 25% 专家时损失约 1%，保留 12.5% 专家时也只掉 3%；同样规模的标准 MoE 在 12.5% 子集下已跌到接近随机水平。\\n\\n更值得注意的是 EMO 真正涌现出的专家语义：聚类后是「健康\u002F医疗」「美国政治选举」「影视音乐」，而不再是「冠词」「所有格」——这意味着我们终于可以让一个稀疏模型按域而不是按 token 去路由，把「加载整模型」换成「按域加载 1\u002F8 专家」成为现实可能。配合 Easy-EP 等现成专家剪枝方法，组合空间相当大。这条路线如果被 DeepSeek、Mixtral 这样的工业级 MoE 采纳，推理侧的显存门槛会再下一个台阶。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fallenai\u002Femo","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"98d29c70-084c-4b55-af54-5947e3e463ea","en","AI2 EMO trains MoE as swappable domain-scheduled modules","Hugging Face user allenai released EMO (Expert Modularity), a method for training MoE (Mixture-of-Experts) models where each expert is a \"pluggable\" module. The standout: 1B-active models can be efficiently customized for specific domains by swapping individual experts, with no retraining.\n\nThe \"pluggable expert\" insight: traditional MoE models have a fixed set of experts, and the entire model must be retrained to specialize for a new domain. EMO's fix: train each expert as a \"pluggable module\" — the expert is trained to be functionally independent, and it can be swapped out without affecting the other experts.\n\nThe \"domain dispatch\" highlight: with EMO, a 1B-active MoE model can be specialized for a new domain by training just a few new experts (not the full model). The new experts are added to the existing expert pool, and the router learns to dispatch to them. The result: a 1B-active model that performs like a domain-specialized 10B model, at the inference cost of 1B.\n\nThe benchmark: on a set of domain-specific tasks (legal, medical, financial), EMO-augmented 1B-active models match the performance of 10B domain-specialized models, at 5× the inference speed. The \"expert swap\" is a one-time cost (training a few new experts), and the runtime cost is unchanged.\n\nThe bigger takeaway: \"modular MoE\" is a significant new direction. The \"train the whole model\" approach is too expensive for domain customization, and the \"expert swap\" approach is significantly more efficient. For the industry, this means \"MoE model marketplace\" will emerge, where developers can pick the right expert modules for their use case. The \"expert library\" becomes a new kind of model IP, and vendors that build the best expert libraries will have a significant advantage.","ai2-emo-emergent-modular-moe-pluggable","2026-05-08T00:00:00Z","2026-06-17T12:19:57.555944Z","2026-08-19T02:08:40.142862Z",true,"agent",102,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"a2b57923-e5f7-4201-8a04-80404796c3e4","开源力量崛起：2026年四月大模型的新平衡","open-source-llm-new-balance-april-2026","2026-04-24T22:03:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"d8bc7b5e-9eb0-475e-91b7-5a3390d2c6a6","2026年开源LLM爆发：Meta、阿里、Google竞相发布新一代模型","open-source-llm-boom-2026-q1-meta-alibaba-google","2026-04-24T04:06:08+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"12801a89-607e-4e90-bada-5e0b66a42e68","Nemotron Cascade 2：NVIDIA万亿参数开源模型的量化革命","nvidia-nemotron-cascade-2-trillion-mixed-precision","2026-04-23T23:05:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00"]