Transformers have a structural blind spot: the compute depth each token receives is fixed—eight layers is eight layers, no matter how long the sequence grows. State tracking, however, demands an update at every input. The Recurrent Looped Transformer (RLT) from Yifan Zhang at Princeton and two University of Pennsylvania collaborators, posted to arXiv on October 6, stitches parallel and recurrent computation into one architecture (arXiv:2610.07591).

Eight layers, split in half

RLT splits its layers between a parallel causal encoder and a recurrent decoder. The encoder processes tokens in parallel and produces representations plus global KV memory; at each token the decoder merges the encoder output with the previous token's final decoder state through a gated merge, backed by per-layer sliding-window attention caches. The computation path grows with sequence length—after t tokens the recurrent path has traversed t times the decoder depth in blocks—while per-token cost stays fixed. The GitHub repo picked up 918 stars within three days, Apache 2.0 licensed.

Trained on 40 bits, 100% at 256

Length generalization tells the story. Trained only on parity of at most 40 bits, the 5+3 and 7+1 splits hit 100% accuracy at 256 bits in every seed, while a same-size eight-layer Transformer sits at 50.07%. On swap-based S5 permutation tracking at eight times the training length (256 operations), the 4+4 split reaches 97.30% versus 0.85%. Modular arithmetic tops out at 93% against 33%. The sixteen-layer series repeats the pattern: 8+8, 9+7, 11+5 and 16+0 all hold 100% at 256 bits; Transformer 16 manages 49.41%.

The feedback path is the whole game

Ablations nail the causal chain: RLT-0, which removes the feedback path, drops parity and S5 back to chance (50.23% at 64 bits) at every split. RLT-2 offers an engineering compromise: update the feedback state once per four-token chunk so known tokens run in parallel—CPU training steps speed up 2.27x and 64-bit parity stays at about 99%, but S5 tracking falls from 100% to about 20%. Permutation tracking genuinely needs per-token feedback. The feedback interval thus becomes a tunable knob: large chunks for pretraining parallelism, then squeeze toward 1 in mid- and post-training.

One execution across training and inference

The repo ships a unified schedule for prefill, generation, pretraining, SFT, and RL replay, with full gradients flowing through recurrent outputs, decoder KV, and encoder memory. A community developer has already reproduced directional results with an independent implementation of roughly 79K parameters: everything fits at training length, but accuracy decays to 60.8% at 128 operations—a reminder that these conclusions cover algorithmic tasks under supervised training, and RL is not evaluated.

The takeaway: this is not another benchmark-chasing paper. It turns depth-versus-state into an explicit design axis, and anyone working on long context or agent memory should read the 108-run comparison in full.