[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-foil-loop-moe-flatten-untie":3,"topics-all":38,"news-related-f1ccbea4-5749-4976-bfbd-9fc835318226":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"f1ccbea4-5749-4976-bfbd-9fc835318226","Foil 重构循环 MoE:专家压进单层,循环次数翻八倍","Foil 把循环 Transformer 与稀疏 MoE 组合:专家参数与单 token 计算量不变,逐次把专家层数减半、每层专家数与循环次数翻倍,再为每次循环配独立注意力。100B token 训练下最扁配置损失比基线低 0.012 nat,解绑注意力再降 0.049 nat,下游平均分提升 3.3 分。","同一个模型、同一份参数预算,只是把专家的排布方式换一下,预训练损失就能再降一截——这是新论文 Foil 交出的答案。它把两条省参数路线——循环 Transformer 和稀疏 MoE——焊在一起,回答了一个没人系统回答过的问题:MoE 到底该怎么循环。\n\n## 两条路线为什么天然互补\n\n循环 Transformer 把一个层块重复使用多次:多花计算、不加参数,把固定规模模型的潜力压榨彻底。稀疏 MoE 则存很多专家、每个 token 只激活少数几个,按需分配计算。\n\n组合起来有个天然好处:每次循环都是一次新的路由决策,token 在不同 pass 里能夠到不同专家组合,专家参数一个不用加。但问题随之而来——参数量和计算量固定下,专家该怎么分布在层和 pass 之间?哪些组件该跨 pass 共享?论文用「压扁 + 解绑」两步作答。\n\n## Foil 的两步手术\n\n第一步是拍扁(flatten):专家参数总量与单 token 专家计算量不变,层数减半、每层专家数与循环次数翻倍——从 8 专家 × 8 层 × 2 次循环,变成 64 专家 × 1 层 × 16 次循环,每次路由都从更大的池子里挑。\n\n第二步是解绑(untie):每次循环配一套独立注意力参数,专家和路由器全局共享。不增加任何计算,却让每个 pass 学到不同的注意力模式。\n\n## 数字说话\n\n20B token 时所有 Foil 配置损失都低于未压扁基线;100B token 时损失随压扁程度单调改善,最扁配置在等参数等计算下比基线低 0.012 nat,下游准确率持平或更好。全扁形状上解绑注意力再降 0.049 nat,三个代表性下游任务平均准确率提升 3.3 分——零额外计算换来的。\n\n## 复现门槛与工程细节\n\n代码以 Apache-2.0 开源在 GitHub(SR-A-W\u002Fhow-to-loop-moe),模型权重同步上架 Hugging Face(ShourenWSR\u002Fhow-to-loop-moe)。训练数据用 FineWeb-Edu 的 sample-100BT 子集(140 个 parquet 分片,约 286GB),分词器用 SmolLM2(词表 49,152)。100B token 的长跑是从 20B run 的 step-45,000 checkpoint 接力续训,而不是从零重来。工程细节:训练和评测必须分环境,torch 分别钉死 2.12.0 和 2.13.0,互不兼容。\n\n论文附了完整 run 对照表:tied 组 S1–S4 与 untied 组 U1–U4 形状一一对应,消融直接可比;模型代码移植自 arXiv:2605.09165 的释出实现,仓库代号 LoopMoE 与 Chen et al.(arXiv:2606.04438)的同名方法无关,后续会改名为 Foil。\n\n## 消融里最值得记住的两条\n\n一是负载均衡不够看。论文发现路由置信度(routing confidence)比负载均衡更能反映专家使用是否健康,其逐 pass 峰值可能是再循环收益递减的信号。二是循环和专家宽度互相放大:每层专家越多,多循环越有用;循环越多,加专家也越有用。合起来就是设计指南:每层专家数和循环次数都往多里堆。\n\n对小模型和边缘部署团队,启发很直接:参数预算不动,重新排布专家层间分布、给每次循环松绑注意力,也能白拿收益。下一步值得盯这套拓扑向更大参数量和产品级模型的迁移——毕竟 0.012 nat 是研究级尺度上量出来的。循环不是免费午餐,但 Foil 至少把菜单排清楚了。\n\n参考:arXiv:2609.35751 · github.com\u002FSR-A-W\u002Fhow-to-loop-moe","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.35751","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"0dff7db7-389d-49a0-8b15-1a2048bde462","en","Foil Flattens Looped MoE Experts and Unties Attention","Foil flattens a looped MoE to 64 experts × 1 layer × 16 passes with untied attention: at 100B tokens loss is 0.012 nat lower; untying adds 3.3 points.","Same model, same parameter budget — just rearrange how the experts are laid out, and pretraining loss drops another notch. That is the answer delivered by Foil, a new preprint that welds together two of today's most popular parameter-saving approaches — looped Transformers and sparse mixture-of-experts (MoE) — and answers a question nobody had systematically addressed before: how exactly should you loop a MoE?\n\n## Why the two roads complement each other\n\nLooped Transformers reuse one block of layers several times: spend extra computation, add no parameters, and squeeze a fixed-size model harder. Sparse MoE takes the other road: store many experts but activate only a few per token, allocating compute on demand.\n\nCombine them and something natural happens: every pass is a fresh routing decision, so a token can reach different expert combinations in different passes without adding a single expert parameter. But a question follows — under a fixed parameter and compute budget, how should experts be distributed across layers and passes, and which components should be shared across passes? The paper answers with two moves: flatten, and untie.\n\n## Foil's two-step surgery\n\nStep one is flattening. With total expert parameters and per-token expert compute held fixed, Foil halves the expert layers, doubles the experts per layer, and doubles the passes. In concrete shape terms, it goes from 8 experts × 8 layers × 2 passes to 64 experts × 1 layer × 16 passes — every routing decision now chooses from a much larger pool.\n\nStep two is untying. Each pass gets its own attention parameters, while experts and routers stay shared. This adds no compute at all, yet lets every pass learn a different attention pattern.\n\n## The numbers\n\nAt 20B tokens, every Foil configuration achieves lower pretraining loss than the unflattened looped baseline. At 100B tokens, loss improves monotonically with the degree of flattening; the most flattened Foil ends 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better. At the fully flattened shape, untying attention lowers loss by another 0.049 nat and improves mean accuracy on three representative downstream tasks by 3.3 points — all with zero additional compute.\n\n## Reproduction cost and engineering details\n\nThe code is open-sourced under Apache-2.0 on GitHub (SR-A-W\u002Fhow-to-loop-moe), with model weights mirrored on Hugging Face (ShourenWSR\u002Fhow-to-loop-moe). Training uses the FineWeb-Edu sample-100BT subset (140 parquet shards, about 286 GB) with the SmolLM2 tokenizer (vocabulary 49,152). The 100B-token runs continue from the step-45,000 checkpoint of a finished 20B run rather than starting over. One engineering wrinkle: training and evaluation require separate environments, pinning torch 2.12.0 and 2.13.0 respectively — mutually incompatible.\n\nThe repository ships a complete run table: the tied group S1–S4 and the untied group U1–U4 share shapes one-to-one, so ablations compare directly. The model code is ported from the released implementation of another looped-language-model work (arXiv:2605.09165); the repo's internal codename LoopMoE is unrelated to the same-named method of Chen et al. (arXiv:2606.04438), and the authors say a later version will rename it to Foil.\n\n## Two ablation findings worth remembering\n\nFirst, load balance alone is not enough. MoE training usually watches load balance, but the paper finds that routing confidence tracks healthy expert use better than load balance does — and the per-pass peak of confidence may signal diminishing returns from further looping. Second, looping and expert width amplify each other: more experts per layer make additional loops more useful, and more loops make additional experts more useful. Together they form a design guide for looped MoE: push both experts-per-layer and passes upward.\n\nFor teams building small models or edge deployments, the takeaway is direct: with the parameter budget untouched, rearranging the layer-wise distribution of experts and untying per-pass attention still yields free gains. What to watch next is whether this topology transfers to larger scales and production-grade models — after all, 0.012 nat was measured at research scale. Looping is not a free lunch, but Foil at least puts the menu in order.\n\nReference: arXiv:2609.35751 · github.com\u002FSR-A-W\u002Fhow-to-loop-moe","foil-loop-moe-flatten-untie","2026-10-07T19:09:13Z","2026-10-07T19:09:15.419818Z","2026-10-07T19:09:15.419829Z",true,"agent",375,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"3559e613-9558-48e1-ab20-f53b62796363","让每个 token 用上全部专家:高德 IntBMoE 解耦参与度、计算与显存,60ms 服务数亿用户","intbmoe-full-participation-block-moe","2026-09-21T13:01:54+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"95e663a4-2cdf-454c-9787-b154fbd41909","TokenRouter:token 级路由提速 64 倍","tokenrouter-token-level-llm-routing","2026-10-09T21:11:28+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"d6ec8624-ad4e-41ce-9566-22d1f0926d49","循环解码器+并行编码器:RLT长度外推翻盘","recurrent-looped-transformer-length-generalization","2026-10-08T23:30:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"f3c43720-650b-4054-9349-a1386e06c8fc","Le Chonk 把法国拉回非美\u002F美头部:38 分的 Mistral Large 4","mistral-large-4-le-chonk-intelligence-index-38","2026-10-08T03:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"44b118b2-cb3d-47bb-8f26-bbce777cfb31","1 亿 rollout 背后:Beam 的 RL 训练工厂","beam-rl-factory-100m-rollouts","2026-10-07T07:15:00+00:00"]