Operand Access and State Dynamics in Looped Transformers
Abstract
Looped Transformers increase computational depth by repeatedly applying shared parameters across recurrent steps, yet it remains unclear how their internal representations support multi-step computation and extrapolation beyond the training horizon. We study modular product learning in looped Transformers, including progressively masked models and Fixed-Point Reasoning Models, using mechanistic interpretability to examine how operand access shapes their computation and extrapolation. We find that the way operands are exposed across recurrent steps, rather than recurrent depth alone, determines how well models extrapolate. Progressive operand exposure extends the compute horizon for commutative composition but fails for non-abelian composition, while among models that learn the task, sequential exposure generalizes farther beyond the training range than causal access. In one successful model, accuracy decreases from 100.0% at sequence length 8 to 10.0% at length 24, with failure emerging gradually through declining readout margins rather than instability or drift in the recurrent state. Representation analyses reveal distinct patterns of information encoding across successful models, suggesting that task success and length generalization alone do not uniquely determine the internal representation of a solution. Together, these results show that operand-access constraints shape recurrent extrapolation without uniquely determining how the underlying computation is represented.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.