Dense attention has a well-known structural waste: after softmax normalization, much of the causal score space receives negligible mass, yet dense GPU kernels still execute the complete post-score path for every QK tile — normalize, multiply by V, write back, no step skipped. The scores are computed, the kernel then discovers "this region barely matters," but the compute is already spent. A new paper from the BAAI team, posted to arXiv on September 26, targets exactly this wasted post-score computation. Titled MassAlloc Attention (MALA), it climbed to the top of Hugging Face Daily Papers within days, accumulating 763 upvotes.
Keep Every QK Score, Allocate the Rest by Mass
MALA's idea splits in two. First, it never prunes QK scoring — every legal causal interaction keeps score access, preserving the integrity of the attention distribution. The second half is the actual contribution: computation after scoring is no longer executed uniformly but allocated by normalized contribution. During the forward pass, MALA borrows its evolving online-softmax normalizer to estimate each region's share of mass in real time; low-contribution regions skip subsequent computation entirely. The backward pass reuses the finalized normalizer to derive nested retained support, using only standard attention state — no auxiliary memory structures. A single tolerance parameter governs both training and inference: how much to skip and how much to keep is decided by the same ruler on both sides.
The Numbers: 3x Backward, 23% Training FLOPs Saved
The paper's numbers are solid. In a matched-work study at 8K context, with total post-score work exactly matched, MALA's mean omitted mass is 0.0188% versus 0.0182% for a per-instance reference-mass oracle — running nearly flush against the theoretical ceiling. On associative-recall tasks, MALA reaches 89.67% accuracy at 8K context against 89.97% for FullAttn, a gap of just 0.3 percentage points. On the engineering side, in an attention-operator benchmark at 128K tokens with tensor parallelism, MALA cuts training forward latency by 2.2x and backward by 3.0x, and speeds up inference decoding by 1.6x; the authors noted on the Hugging Face paper page that the test environment was 8xH100 with TP=8. Scale validation matters more: across scaling-law training from 0.6B to 14B parameters, MALA tracks FullAttn in perplexity while reducing total training FLOPs; the 14B model training at 32K context saves 23.1% of total FLOPs, and a separately continued-trained 32B model achieves comparable knowledge, reasoning, and long-context retrieval scores to FullAttn.
Not the Same Road as Sparse Attention
The lane is crowded. LEMA, also from September, moves KV cache out of VRAM; HyQuant does mixed-precision attention quantization; Video DeltaNet swapped hybrid attention into video generation — all chewing on the same question of how the attention compute budget should be spent. MALA's differentiator is that it presupposes no sparse structure. Fixed-window and top-k routes require humans to tell the model in advance where sparsity roughly lives; MALA hands the decision back to the softmax statistics themselves — the distribution already tells you what matters, so allocate accordingly. Another easily missed detail: the backward speedup (3.0x) exceeds the forward one (2.2x). This is fundamentally a training-floor algorithm, and for labs whose budgets are dominated by pretraining, 23% of FLOPs translates directly into real money.
Pour Some Cold Water
The paper page shows "Models citing this paper: 0," and the Hugging Face page has no code repository link — third-party replication is, for now, impossible, and every benchmark number is self-reported. The sensitivity of the tolerance hyperparameter and cross-architecture transferability must wait for artifacts to land. The same team also posted a sibling paper, CoWindow Attention (arXiv:2609.32704), which lets multiple attention heads share the work of accessing history and saves 28.5% of FLOPs at 14B training — reading the two together reveals BAAI's full layout on attention compute allocation.
The authority to allocate attention's compute budget is migrating from fixed rules set by humans to the distribution's own statistics. If MALA's tolerance mechanism lands in mainstream training frameworks, the cost curve of long-context training will bend once more — and the next 128K-scale pretraining bill may come back a slice smaller.
References: arXiv:2609.32712 (https://arxiv.org/abs/2609.32712) · Hugging Face paper page (https://huggingface.co/papers/2609.32712)