Learning to Loop: Reinforcement Learning for Looped Transformers
Abstract
Looped Transformers can increase computation by repeatedly applying existing layers, but most approaches use fixed repetition patterns across inputs. We ask whether a model can instead learn, for each question, when additional computation is useful and where it should occur. We introduce a lightweight controller that observes the model's hidden states and decides after each layer whether to repeat the layer or proceed forward. The controller is trained by comparing alternative computation routes for the same question using signals from final-answer correctness and the model's likelihood of a reference solution, while the language model itself remains frozen. Unlike fixed-depth looping, our method allows different questions to use different numbers and locations of repetitions under a maximum computation budget, while retaining fixed-depth strategies as special cases. We evaluate the approach across three 7–8B language models and four mathematical reasoning benchmarks: GSM8K and MathQA for in-domain evaluation, and SVAMP and MATH-500 for frozen transfer. Adaptive routing improves reasoning performance over fixed repetition schedules while allocating different amounts and locations of computation across questions. Controlled comparisons with exact-depth routing and shuffled question–route assignments further indicate that the gains arise from learning question-specific computation paths rather than simply executing more layers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.