[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-flexmoe-qwen2-57b-elastic-subnet":3,"news-related-15e5549b-b18a-43fd-a500-6008bc2709d7":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"15e5549b-b18a-43fd-a500-6008bc2709d7","FlexMoE 把 MoE 大模型压成「弹性子网络族」：一次训练多档压缩，Qwen2-57B 剪掉 50% 专家仍保 99.8% 性能","MoE 架构虽以「稀疏激活」著称，但所有专家仍要常驻显存——这让大模型部署成本居高不下。arXiv 2606.27866 上的 FlexMoE 提出一种「一次性训出全套预算」的思路：先把每个专家的 FFN 通道按重要性排序，让专家各自学一个离散动作剪掉低权重通道，再以渐进加压从同一个训练 run 里导出从高到低多档预算下的子网络。换句话说，一次训练就能拿到一个「可按预算弹性拉伸」的嵌套子网络族。\n\n更值得称道的是它的「跨预算迁移」设计：在中等预算（40%）上做一次恢复式微调，恢复后的模型可直接迁移到其他未见预算档位，无需重新训练。论文在 Qwen2-57B-A14B 上展示了惊人的保真度——无微调剪掉 50% 路由专家参数时仍可保留 99.8% 的基座性能；剪得更多时，部署侧能拿到真实的显存下降和吞吐增益，并支持运行时在线切换预算，无需为不同 SLA 各压一份权重。\n\nFlexMoE 把「嵌套结构 + 一次训练多档输出」摆到了 MoE 大模型面前：推理服务方只需一份 FlexMoE 化的权重，就能在低配边缘环境和高吞吐数据中心之间无缝切换。作者把「kernel 级 co-design」和「online budget switching」放在最后，正是为了告诉产业——MoE 部署第一次具备「按预算弹性伸缩」的工程能力，这是相对 Mistral \u002F DeepSeek \u002F Qwen 等 MoE 大模型都能直接落地的实用主义工具。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.27866","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"aea1f85c-fbcb-4688-961b-b07e9e1720a9","en","FlexMoE: one training, many compression tiers, 99.8% intact","Although MoE architectures are celebrated for sparse activation, all experts still need to reside in VRAM — and that keeps deployment costs stubbornly high. arXiv 2606.27866 introduces FlexMoE, a \"train-once-and-get-the-whole-budget-stack\" approach: each expert's FFN channels are sorted by importance, every expert learns a discrete action that prunes its low-weight channels, and progressive compression then exports a nested sub-network family from high to low budgets out of a single training run. In other words, one run delivers a \"budget-elastically stretchable\" nested family of sub-networks.\n\nMore impressive is its \"cross-budget transfer\" design: a single recovery fine-tune at the medium 40% budget lets the recovered model transfer directly to unseen budget tiers, with no retraining. On Qwen2-57B-A14B the paper shows striking fidelity — without fine-tuning, dropping 50% of routed expert parameters still preserves 99.8% of base performance; pushing further down yields real VRAM savings and throughput gains at deployment, with runtime online budget switching that removes the need to compress a separate weight for each SLA.\n\nFlexMoE puts \"nested structure + multi-tier output from one training run\" squarely in front of MoE LLMs: inference providers need only one FlexMoE-style checkpoint to switch seamlessly between low-resource edge and high-throughput data-center environments. The authors place \"kernel-level co-design\" and \"online budget switching\" at the end on purpose — to tell the industry that MoE deployment now has \"elastic budget scaling\" as an engineering capability, a pragmatic tool any MoE model from Mistral \u002F DeepSeek \u002F Qwen can adopt directly.","flexmoe-qwen2-57b-elastic-subnet","2026-06-29T14:00:00Z","2026-06-29T14:09:34.544026Z","2026-08-19T02:08:40.142862Z",true,"agent",120,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b0c434a5-4911-4297-b1ef-2c44cbc26653","蚂蚁新研究:19769 个代码仓库,炼出百万条 agent 技能","code2skill-agent-skill-synthesis","2026-09-21T19:06:31+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"b0c4e8d2-5662-4e3e-b489-6202eabbe97b","Dream-RSI 把历史当模拟器:162 倍杠杆重写 RSI 算力账本","dream-rsi-replay-simulator-162x","2026-09-16T06:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2266cea6-06f1-4932-8905-1bc3f2e5a8c0","Meta FAIR 字节蒸馏研究:End-Of-Token 渐近反超 token 蒸馏 4%,数据只需 1\u002F6","meta-fair-byte-distillation-token-ceiling-2026-09","2026-09-15T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e91b3add-4f1d-48d3-ab7a-bd2c6e8c1765","QuIP 崩、OPTQ 降级:Kashin-DCT 在 4-bit 量化压力测试里活了下来","kashin-dct-2bit-llm-quantization","2026-09-12T15:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c83af54b-79ed-445c-9482-07d98c26c36b","BeaconKV:长推理会回头看,只压最近窗口的 KV 缓存注定丢东西","beaconkv-beacon-query-kv-cache-compression","2026-09-09T11:25:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"21c7dec1-f68e-4641-974c-ae2bce87393e","教师打分、验证器掌舵:腾讯混元 FlowBalance 给自蒸馏装上方向门控,Qwen3-8B 数学均值超 GRPO 2.12 分","flowbalance-verifier-gated-self-distillation","2026-09-08T15:08:17+00:00"]