For the past decade, from the original Transformer to Gated Attention, Hope-attention, and Titans, almost all modern language models share a default skeleton: L layers of the same-structure block stacked uniformly, with each layer getting an equal share of parameters. This design inherited from the original transformer has hardly ever been questioned.

Reza Bayat et al. at Mila, Cornell, Université de Montréal, and the CIFAR AI Chair challenge it in arXiv:2606.23670. They ran a simple controlled experiment: on a 440M-parameter transformer, holding total parameters constant, they simply redistribute the MLP intermediate width across three segments (front / middle / back). The result is strikingly asymmetric — "wide front, narrow back" lowers perplexity by 0.32 versus the uniform baseline, while the reverse "narrow front, wide back" runs more than a full point higher. Get the direction wrong and you waste budget; get it right and the gain is free.

Based on this observation, they propose Tapered Language Models (TLMs): under fixed parameter and FLOPs budgets, use a smooth cosine schedule to monotonically decrease the MLP width along depth, front-loading capacity to the shallow layers. The design space has three schedules (linear / cosine / sigmoid); cosine is the most stable because it has plateaus at both ends and the smoothest transition.

The core result is hard: TLMs stably lower perplexity and improve downstream benchmarks on four token-mixing-different architectures (Transformer, Gated Attention, Hope-attention, Titans) and three scales (440M / 760M / 1.3B) — with no extra parameters or compute. The mechanism is also explained: the authors measure the alignment between each layer's MLP output and the residual stream, and find that the deeper the layer, the more it tends to "restate" the residual rather than "write" new features — so spending width on such redundant layers is wasteful.

Layer-skipping, ShortGPT, early exit, and model pruning have all repeatedly hinted that "deep MLPs aren't that important." TLMs push that thread from "can we cut" to "in what proportion should we allocate." It needs no new compute and no new data — a free lever any team can pick up by changing the initialization schedule. The next step worth watching is whether this "non-uniform along depth" thinking spreads to attention-head counts, KV dimensions, recurrent state sizes, and even the number of MoE experts — all of which also have a "front and back layers contribute unevenly" smell.