acceptodds
Under review as a conference paper at ICLR 2027

How to Loop Your Transformer: Loop Scaling, Specialization, and Parallel Inference

Abstract

Looping in transformer models reuses weights by repeatedly applying the same transformer, or parts of it, to token representations, effectively making the model deeper without increasing its parameter count. Whether looping is useful depends on which constraint dominates: model size, training compute, or autoregressive inference efficiency. We study these trade-offs jointly in dense decoder-only transformers trained from scratch, varying which parts of the model are looped, how many loops are applied, and how information is passed between loops. Our experiments span models with up to 1B parameters, trained on up to 155B tokens using up to training FLOPs, with up to loops. We find that looping the full transformer generally gives the strongest quality at a fixed parameter budget, whereas looping a middle block of layers gives a better return when training FLOPs are constrained. These trade-offs also depend on the number of loops: adding loops is not uniformly beneficial, and gains eventually saturate, reverse, or become unstable. This saturation motivates asking whether forcing every loop to use exactly the same weights limits the benefit of deeper looping. We address this by adding small loop-specific low-rank updates that improve quality, with larger gains as more loops are added. A separate practical concern for looped transformers is the sequential dependency during autoregressive generation, which causes a slowdown proportional to the number of loops. Parallel Loop Transformers address this limitation by passing information between loops across adjacent token positions via hidden states, allowing loop stages to run concurrently during decoding. However, in our experiments the original formulation does not consistently retain the quality gains of ordinary looping. We find that preserving the current-token embedding and separately normalizing it and the previous-token hidden state before combining them recovers strong performance across model scales and training budgets. Together, these results show that effective looping requires matching the design to the resource constraint and inference setting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.

How to Loop Your Transformer: Loop Scaling, Specialization, and Parallel Inference | acceptodds