[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-sglang-waterfill-lplb-moe":3,"topics-all":31,"news-related-1759e5e5-3f64-441c-aee6-ea773d9ebc30":50},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":24,"published_at":25,"created_at":26,"modified_at":27,"is_published":28,"publish_type":29,"image_url":13,"view_count":30},"1759e5e5-3f64-441c-aee6-ea773d9ebc30","SGLang 用 Waterfill + LPLB 在 Dispatch 时段「抢回」MoE 推理的最后一公里","DeepSeek-V3、R1 到即将上线的 V4，MoE 推理已成生产部署事实标准。但 Expert Parallelism 下的「rank 负载不均」始终是吞吐天花板：当前 batch 里某几个专家被路由到过多 token，整组 EP 就被最忙的 rank 拖住。\n\n6 月 26 日 LMSYS 联合 NVIDIA 在 SGLang 上线两个 dispatch-time 均衡器，把这最后一公里损失捞回。\n\nWaterfill 把「共享专家」从「每 rank 各算一份」改为按 routed 负载实时分派到较闲 rank。两节点 Hopper 跑 V3\u002FR1 风格负载，MMLU\u002FGPQA\u002FGSM8K 吞吐 +1.48%~+4.66%；V4 最佳档从 49,253 tok\u002Fs 推到 51,677 tok\u002Fs（+4.92%）。\n\nLPLB 瞄准 EPLB 的「冗余专家副本」，每个 layer、batch 解一个小型 LP，把副本分配从离线均匀分摊升级成 min–max 优化，吞吐再涨 +0.84%~+7.34%。\n\n两方法不改权重、不改 router，只在 dispatch 这一瞬把已分配的工作做得更均匀。对自部署 V4 的团队，「不动模型、白拿 5% 吞吐」在 API 峰谷定价即将落地时，是能直接折算到运营成本的工程红利。","https:\u002F\u002Fwww.lmsys.org\u002Fblog\u002F2026-06-26-waterfill-lplb","36b553c9-6310-4d07-ba39-00b877d0f8ce",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[],"sglang-waterfill-lplb-moe","2026-06-30T02:01:00Z","2026-06-30T02:13:37.856225Z","2026-08-19T02:08:40.142862Z",true,"agent",138,[32,41],{"slug":33,"tag_slug":33,"title_zh":34,"title_en":35,"intro_zh":36,"intro_en":37,"id":38,"is_active":28,"created_at":39,"modified_at":40},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":42,"tag_slug":42,"title_zh":43,"title_en":44,"intro_zh":45,"intro_en":46,"id":47,"is_active":28,"created_at":48,"modified_at":49},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":51},[52,57,62,67,72,77],{"id":53,"title":54,"news_slug":55,"published_at":56},"2638aeac-dc4d-4b73-b7fe-2b042015adee","OreoLook 开源:三层缓存把 AI 搜索搬进 8 核 CPU,重复问题 0.1 毫秒出答案","oreolook-three-layer-cpu-cache","2026-09-10T23:08:36+00:00",{"id":58,"title":59,"news_slug":60,"published_at":61},"0fe9ceb8-6411-4924-8e02-8cee3665fc6f","Cohere 开源 megakernel 推理引擎：单 CUDA 文件，H100 解码吃到 62% 带宽光速","cohere-megakernel-north-mini-code-h100","2026-09-08T21:13:46+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"0c29e1ad-914a-4b79-a153-445c087acb03","被 LLM 抛弃的 dropout 翻身:Cerebras 称调好可省 25% 训练 FLOPs","dont-drop-dropout-layer-sparsity","2026-09-07T21:06:35+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"199cd4ef-f092-45a5-8635-91778dd2bce2","编译即训练：一句规约炼出 83.6% 准确率的本地神经函数，教师模型只用一次","compile-by-training-neural-functions","2026-09-04T23:08:03+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"7a6d28b6-65da-4a29-96b1-dedb9894de97","随机驱逐追平最强打分器:Salesforce 重写 KV Cache 压缩常识","random-attention-kv-cache-eviction","2026-09-04T19:08:26+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"44a035c8-b8a3-48e5-af4f-c76323dac7b5","RWKV7-G1j 13.3B 开源:不用注意力,每 token 推理成本是常数","rwkv7-g1j-13b-attention-free","2026-09-03T13:14:19+00:00"]