[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-maestro-moe-expert-pruning-markov":3,"topics-all":36,"news-related-b2625716-65d0-4ec2-a6b2-6a6addb67721":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b2625716-65d0-4ec2-a6b2-6a6addb67721","MAESTRO 把 MoE 专家剪枝扔进马尔可夫链：50% 压缩下鲁棒性维度反涨 10pp","arXiv 2607.08601 提出的 MAESTRO 把 MoE 专家剪枝从「局部打分」推到「全局 Markov 链平稳分布」。每个 token 只激活少量参数,但所有专家权重都得常驻显存——MoE 大模型部署这层「内存税」被诟病已久。MAESTRO (Markov-chain Approximated Expert Sparsification via Transition-based ROuting) 没有去卷新架构,而是盯住最朴素的工程问题:哪些专家值得留？\n\n作者把自回归的专家激活轨迹建模成遍历马尔可夫链,让稳态分布编码跨层依赖,作为全局重要性启发函数。常见 layer-wise 打分方法只盯着 token 在当前层喜欢谁,MAESTRO 关心的是专家之间长期的「对话模式」:谁被反复连叫,谁只是偶发路过,稳态概率一字排开。\n\n50% 严格压缩下,跨 Safety、Bias、Ethics 等五个领域平均性能保留比 SOTA baseline 高出最多 10.61 个百分点,跨任务方差也显著降低。换句话说,剪掉的不只是分数低的专家,而是「破坏路由一致性」的那批——传统启发容易把「偶尔被叫到的备用专家」误判为冗余,留下真正干扰下游特征的「路由噪声节点」。\n\n最有意思的不是 50% 压缩本身,而是它在鲁棒性维度上的稳定性。MoE 裁剪最怕「能力一起裁没了,偏见悄悄留下来」——传统 layer-wise 打分做不到的事,MAESTRO 用平稳分布给出的全局剪枝启发给出了一个实证答案。对所有还在纠结「如何把 200B+ MoE 实际部署」的团队,这是一份把\"剪枝\"从经验手艺升级到带数学基底方法的清晰示范。","https:\u002F\u002Farxiv.org\u002Fhtml\u002F2607.08601v1","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"474f13ac-4d95-4d34-809e-baa30de50f11","en","MAESTRO prunes MoE experts via Markov chains, +10pp robust","arXiv 2607.08601's proposed MAESTRO pushes MoE expert pruning from \"local scoring\" to \"global Markov-chain stationary distribution\". Each token only activates a few parameters, but all expert weights must reside in memory — the \"memory tax\" of deploying MoE large models has long been criticized. MAESTRO (Markov-chain Approximated Expert Sparsification via Transition-based ROuting) doesn't go for new architectures, but focuses on the most basic engineering question: which experts are worth keeping? The authors model the autoregressive expert activation trajectory as an ergodic Markov chain, letting the stationary distribution encode cross-layer dependencies as a global importance heuristic. Common layer-wise scoring methods only look at which experts the token likes in the current layer; MAESTRO is concerned with the long-term \"conversation pattern\" between experts: who gets called repeatedly, who is just a passing visitor, and the stationary probability lays it all out. Under 50% strict compression, average performance retention across five domains (Safety, Bias, Ethics, etc.) outperforms the SOTA baseline by up to 10.61 percentage points, and cross-task variance also drops significantly. In other words, what's pruned isn't just the low-scoring experts, but the ones that \"break routing consistency\" — traditional heuristics easily misjudge the \"occasionally called backup expert\" as redundant, leaving behind truly interfering \"routing-noise nodes\". What's most interesting isn't the 50% compression itself, but the stability on the robustness dimension. MoE pruning's biggest fear is \"capability pruned together, bias quietly remains\" — what traditional layer-wise scoring can't do, MAESTRO's global-pruning heuristic given by stationary distribution provides an empirical answer. For all the teams still struggling with \"how to actually deploy 200B+ MoE\", this is a clear demonstration of upgrading \"pruning\" from empirical craft to mathematically-grounded method.","maestro-moe-expert-pruning-markov","2026-07-09T15:32:54Z","2026-07-11T20:11:36.331596Z","2026-08-19T02:08:40.142862Z",true,"agent",175,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"6df953d8-5371-47e5-94e1-2a4a0d629e4a","训练时压缩SSM：MIT CompreSSM如何让状态空间模型「边学边瘦」","mit-compressm-ssm-training-time-compress","2026-06-02T13:15:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"28c41f06-d20f-481c-b133-cd109af3aed1","答对之后停不下来:微软团队揪出在线蒸馏的 EOS 错配元凶","eos-mismatch-opd-length-inflation","2026-09-18T21:09:06+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"e91b3add-4f1d-48d3-ab7a-bd2c6e8c1765","QuIP 崩、OPTQ 降级:Kashin-DCT 在 4-bit 量化压力测试里活了下来","kashin-dct-2bit-llm-quantization","2026-09-12T15:10:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"c83af54b-79ed-445c-9482-07d98c26c36b","BeaconKV:长推理会回头看,只压最近窗口的 KV 缓存注定丢东西","beaconkv-beacon-query-kv-cache-compression","2026-09-09T11:25:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"9825e20d-c9eb-4300-99b5-12eb7d0e755d","自信的错误教师最危险:TGOPD 给在线蒸馏装提示级门控,教师 GPU 利用率 9.8% 升至 78.9%","tgopd-teacher-gated-on-policy-distillation","2026-09-08T23:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00"]