Start with an uncomfortable finding: the pretrained LLMs you use every day likely cannot follow even a short chain of references like K = apple; B = K; D = B; print(D) when asked to answer directly. A Georgia Tech team — Zehao Jin, Ruixuan Deng and Junran Wang — measured this in a paper submitted to arXiv on September 29 (2609.36585): thirteen base models from 0.6B to 32B parameters reliably follow only 1.4 to 3.6 lines of such chains when no chain-of-thought is allowed. More uncomfortably, adding extra pretrained loops barely helps. The depth is sitting right there; the model just does not use it by default.
One LoRA, everything else frozen
The fix is almost austere: attach a single rank-8 LoRA at one early layer, freeze every other weight, and train only that adapter. The effect is an order-of-magnitude jump — Qwen3-8B goes from 15.5% to 99% exact accuracy on 24-line chains, and a longer-trained LoRA reaches 50 lines. The 1.4B Ouro model reaches 60 lines after four loops and at least 160 after eight. The team open-sourced the full reproduction pipeline (Lunamos/stop-thinking-too-early, Apache 2.0); the headline LoRA (Qwen3-8B, layer 14) trains and evaluates in about 16 minutes on a single A100 — a price any inference-optimization team can afford.
The relay mechanism: the computation was already there
Why can a one-layer adapter mobilize dozens of frozen layers? The paper's analysis offers an elegant answer: the LoRA starts a relay. Program lines pass their chain identity down through a short range of middle layers, and frozen attention heads read progressively further up the chain; remove attention to the parent line and the relay stops immediately. This confirms the paper's title — transformers do not lack the ability to compute, they simply stop thinking too early: the default forward pass uses only a small slice of the available depth. The team also demonstrates a frozen-model measurement that locates the last useful intervention layer within tolerance in three of four held-out models, and task-specific LoRAs also improve results on MuSiQue — the mechanism is not confined to synthetic tasks.
Three takeaways for the industry
First, "the model cannot do it" and "the model was never activated" are different failure modes. When an eval returns a low score, do not rush to buy a bigger model — it may be a depth-utilization problem that a 16-minute LoRA can flip. Second, the inference-cost narrative needs revision: if a tiny parameter edit releases computation already latent in a frozen network, there is a cheap middle road between "buy more depth" and "activate what you have". Third, interpretability gains an engineering handle: the relay mechanism maps a clean division of labor across layers, which anyone working on inference acceleration or architecture search can build on.
So the next time a model flubs a long-chain question, do not blame capacity too quickly — it may simply not have been woken up. Paper, code and an interactive demo are linked below; 16 minutes on one A100 is worth spending yourself.
References: arxiv.org/abs/2609.36585 · github.com/Lunamos/stop-thinking-too-early