[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-iclr-2026-capacity-aware-moe":3,"news-related-f9a6c57b-a98e-407c-9751-93907ff314b0":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"f9a6c57b-a98e-407c-9751-93907ff314b0","Capacity-Aware Inference：MoE 推理 1.85× 加速","MoE 想做推理，最怕的不是模型大,而是专家路由不均——少数专家被挤爆,其他专家干等,整个 batch 被「最慢那个」拖住,这在论文里被叫做 Straggler Effect。UMD 的 Shwai He 等人把这现象形式化,并在 ICLR 2026 上提出 Capacity-Aware Inference:先用 Capacity-Aware Token Drop 给每个专家设一个容量上限,把超载专家的溢出 token 直接丢掉,换来最多 30% 的加速(OLMoE 上只掉 0.9 个点);再升级到 Expanded Drop,在丢之前先把 token 路由到同卡上负载更低的备选专家,在 Mixtral-8x7B-Instruct 上跑出 1.85× 推理加速的同时平均还涨了 0.2 个点。最关键的一点是:这套方法是纯 inference-time 的,不动权重、不重训,直接用 apply_capacity_aware_moe_patch 就能套到现有 MoE checkpoint 上,对 OpenMoE \u002F DeepSeek-V3 \u002F Mixtral 这类已经在生产里跑的稀疏模型,等于零成本薅羊毛。代码已开源(case-lab-umd\u002FCapacity-Aware-MoE,star 20,90 commits),同时打通了 lm-evaluation-harness 和 VLMEvalKit 两条评估链路,从纯文本到多模态 MoE 都能验证。对工程团队的启示是明确的:在做 MoE 推理优化时,「专家路由不均」往往比「专家激活太稀疏」更值得优化——前者是木桶的短板,直接决定 P99 延迟。论文 arXiv:2503.05066v5,ICLR 2026 接收。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2503.05066","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"03ab2c27-e583-4152-8f31-33412d6a0e0f","en","Capacity-Aware Inference: MoE inference 1.85x faster","For MoE inference, the worst enemy isn't model size — it's uneven expert routing: a few experts get crushed, the others sit idle, and the whole batch is dragged down by \"the slowest one\", a phenomenon the paper calls the Straggler Effect. Shwai He et al. from UMD formalized this and, at ICLR 2026, proposed Capacity-Aware Inference: first, Capacity-Aware Token Drop sets a per-expert capacity ceiling, dropping overflow tokens from overloaded experts for up to 30% speedup (only a 0.9-point drop on OLMoE); then Expanded Drop, before dropping, reroutes tokens to a less-loaded alternate expert on the same card, achieving 1.85x inference speedup on Mixtral-8x7B-Instruct while actually gaining 0.2 points on average. Crucially, this is a purely inference-time method: it doesn't touch the weights, doesn't retrain, and can be dropped onto any existing MoE checkpoint with `apply_capacity_aware_moe_patch` — so for sparse models already in production like OpenMoE \u002F DeepSeek-V3 \u002F Mixtral, it's free money. The code is open-sourced (case-lab-umd\u002FCapacity-Aware-MoE, 20 stars, 90 commits), and it has been wired into both lm-evaluation-harness and VLMEvalKit, so it can be validated on pure-text and multimodal MoE alike. The lesson for engineering teams is clear: when optimizing MoE inference, \"uneven expert routing\" is usually a bigger lever than \"experts are too sparsely activated\" — the former is the wooden barrel's short plank and directly determines P99 latency. arXiv:2503.05066v5, accepted at ICLR 2026.","iclr-2026-capacity-aware-moe","2026-07-27T10:00:00Z","2026-07-27T10:04:53.277449Z","2026-08-19T02:08:40.142862Z",true,"agent",148,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"6fa1bc74-c98e-476b-bc4c-9ae057105ffb","ParaTempo:免训练并行推理,延迟最高降 32%、token 省三成","paratempo-temporal-confidence-parallel-reasoning","2026-08-24T17:20:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"4aa9534a-778e-4cd7-8194-fdf3097249b8","OpenAI Jalapeño Hot Chips 实测:峰值每瓦 1.9×,延迟压到 1 秒","openai-jalapeno-hot-chips-benchmark-2026","2026-08-26T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"259d91b2-ed6b-4af8-8f2a-f759b84cc617","蚂蚁 Ling-3.0 Flash：124B\u002F5.1B MoE 的 Agent 生产级模型","inclusionai-ling-3-flash-hybrid-linear-moe-agent","2026-08-14T08:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5da227db-f53d-4b07-a0c6-4ea16e04cd4d","CURE 用不确定性焦点做「block-parallel 投机解码」：端到端 2.66–3.49×、接受长度涨 4.2–7.5%","cure-block-parallel-speculative-decoding","2026-08-08T02:00:00+00:00"]