Same model, same parameter budget — just rearrange how the experts are laid out, and pretraining loss drops another notch. That is the answer delivered by Foil, a new preprint that welds together two of today's most popular parameter-saving approaches — looped Transformers and sparse mixture-of-experts (MoE) — and answers a question nobody had systematically addressed before: how exactly should you loop a MoE?

Why the two roads complement each other

Looped Transformers reuse one block of layers several times: spend extra computation, add no parameters, and squeeze a fixed-size model harder. Sparse MoE takes the other road: store many experts but activate only a few per token, allocating compute on demand.

Combine them and something natural happens: every pass is a fresh routing decision, so a token can reach different expert combinations in different passes without adding a single expert parameter. But a question follows — under a fixed parameter and compute budget, how should experts be distributed across layers and passes, and which components should be shared across passes? The paper answers with two moves: flatten, and untie.

Foil's two-step surgery

Step one is flattening. With total expert parameters and per-token expert compute held fixed, Foil halves the expert layers, doubles the experts per layer, and doubles the passes. In concrete shape terms, it goes from 8 experts × 8 layers × 2 passes to 64 experts × 1 layer × 16 passes — every routing decision now chooses from a much larger pool.

Step two is untying. Each pass gets its own attention parameters, while experts and routers stay shared. This adds no compute at all, yet lets every pass learn a different attention pattern.

The numbers

At 20B tokens, every Foil configuration achieves lower pretraining loss than the unflattened looped baseline. At 100B tokens, loss improves monotonically with the degree of flattening; the most flattened Foil ends 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better. At the fully flattened shape, untying attention lowers loss by another 0.049 nat and improves mean accuracy on three representative downstream tasks by 3.3 points — all with zero additional compute.

Reproduction cost and engineering details

The code is open-sourced under Apache-2.0 on GitHub (SR-A-W/how-to-loop-moe), with model weights mirrored on Hugging Face (ShourenWSR/how-to-loop-moe). Training uses the FineWeb-Edu sample-100BT subset (140 parquet shards, about 286 GB) with the SmolLM2 tokenizer (vocabulary 49,152). The 100B-token runs continue from the step-45,000 checkpoint of a finished 20B run rather than starting over. One engineering wrinkle: training and evaluation require separate environments, pinning torch 2.12.0 and 2.13.0 respectively — mutually incompatible.

The repository ships a complete run table: the tied group S1–S4 and the untied group U1–U4 share shapes one-to-one, so ablations compare directly. The model code is ported from the released implementation of another looped-language-model work (arXiv:2605.09165); the repo's internal codename LoopMoE is unrelated to the same-named method of Chen et al. (arXiv:2606.04438), and the authors say a later version will rename it to Foil.

Two ablation findings worth remembering

First, load balance alone is not enough. MoE training usually watches load balance, but the paper finds that routing confidence tracks healthy expert use better than load balance does — and the per-pass peak of confidence may signal diminishing returns from further looping. Second, looping and expert width amplify each other: more experts per layer make additional loops more useful, and more loops make additional experts more useful. Together they form a design guide for looped MoE: push both experts-per-layer and passes upward.

For teams building small models or edge deployments, the takeaway is direct: with the parameter budget untouched, rearranging the layer-wise distribution of experts and untying per-pass attention still yields free gains. What to watch next is whether this topology transfers to larger scales and production-grade models — after all, 0.012 nat was measured at research scale. Looping is not a free lunch, but Foil at least puts the menu in order.

Reference: arXiv:2609.35751 · github.com/SR-A-W/how-to-loop-moe