Simply Stabilizing the Loop via Fully Looped Transformer
Abstract
Scaling model performance typically requires increasing model size. Looped Transformer (LT) offers a parameter-efficient alternative by repeatedly reusing the same Transformer blocks, converting additional computation into performance gains without increasing the parameter count or context length. However, increasing the number of loop iterations makes LT increasingly difficult to optimize, often leading to gradient oscillation, residual explosion, and training collapse. We propose the Fully Looped Transformer (FLT), a simple, parameter-free redesign that addresses these problems without adding learnable parameters or imposing specially designed optimization constraints. FLT combines a Fully Looped Architecture, which exposes the recurrent state to every layer, with Attention Injection, which controls recurrent information flow through the existing attention module. These modifications enable FLT to train stably through 12 loops in our tested configurations, while baseline looped models enter high-loss regimes. In the Base-size, 6-loop setting, FLT improves downstream-task accuracy by 14.9% relative to the original LT. Fixed-loop checkpoint evaluations further show that performance improves as inference depth approaches the trained depth, even without randomized loop-count training, revealing a depth-aware test-time compute interface. Extending this interface beyond the trained range is a promising research direction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.