acceptodds
Under review as a conference paper at ICLR 2027

LT2: Linear-Time Looped Transformers

Abstract

Looped Transformers (LT) reuse the same layers several times before producing each token, gaining computational depth without adding parameters. However, this *parameter efficiency* does not directly translate into *computational efficiency*: with full softmax attention, the quadratic cost is incurred at every loop, making both training and inference slow. We therefore study **LT2 (Linear-Time Looped Transformers)**, a family of looped architectures that replace quadratic softmax attention with linear-time attention, enabling fast looping over a fixed-size recurrent memory. We first show, both theoretically and empirically, that looping and linear attention are strongly synergistic: on recall—linear attention's known weakness—looping helps linear attention more than it helps full attention, improving natural-text recall by +8.9 points versus +4.8 for full attention. We then propose **historical KV replay**, a simple, parameter-free mechanism that co-designs *recurrence in depth* (looping) and *recurrence in time* (linear attention): at every loop, each token replays its key–value pairs from all earlier loops, in reverse order, into the current loop's state, so that each update draws on historical information as well as the current loop. Applied to a vanilla looped Gated DeltaNet (GDN), historical KV replay improves natural-text recall from 34.6 to 41.1, strengthens needle-retrieval extrapolation, and raises performance on every downstream benchmark, by 2.7 points on average. Building on these insights, we further propose a cheaper variant that replays only the first loop's key–value pair instead of the full history. It runs at the cost of plain looping yet recovers about 70% of the full replay's gain, offering the best efficiency–quality trade-off in practice. Finally, we provide a simple recipe for converting a pretrained looped Transformer into a hybrid LT2. With only 1.3B tokens of continued training, our *Ouro-LT2-1.4B* retains most of its teacher's quality, letting practitioners reuse existing checkpoints rather than train from scratch.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.