Loop-μP: Residual Scaling and Hyperparameter Transfer in Looped Transformers
Abstract
Looped Transformers have drawn growing attention because looping improves multi-step reasoning and increases the executed depth of a model without adding parameters. They achieve this by reusing a stack of layers for loops. However, increasing the loop count not only destabilizes training but also shifts the best hyperparameters, so they must be searched again for every loop count. Inspired by Depth-P (Yang et al., 2024a), which multiplies each residual branch by for depth so that a learning rate tuned on a shallow untied network transfers to a deep one, we study how the residual branches and the learning rate should scale with the loop count . We find that loops cannot simply be treated as extra depth, because the loops share weights. At initialization, distinct layers add nearly uncorrelated increments to the residual stream, so their sum grows as , whereas a shared layer adds positively correlated increments at every loop, so their sum grows linearly in . For a single layer looped times, we show in the infinite-width limit that multiplying the residual branch by and keeping the learning rate independent of is the only parameterization with both a stable forward pass and maximal residual-stream updates. We then generalize the analysis to a stack of distinct layers looped times to obtain a parameterization we call Loop-P, in which the loop factor combines with the depth factor . Empirically, we find in language modeling that the base learning rate tuned on the smallest model stays best under Loop-P for loop counts up to and executed depths up to layers. These results make Loop-P a practical recipe for scaling looped Transformers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.