What Does Another Loop Buy? Attention Reallocation Tracks the Marginal Value of Test-Time Compute in Looped Transformers
Abstract
Looped transformers repeatedly apply shared weights, making iteration count a natural control for test-time compute. However, the benefit of additional iterations is non-monotonic, while existing adaptive stopping methods typically require training or rely on output-level heuristics. We show that the marginal value of a loop is reflected in attention reallocation between successive iterations. Across controlled synthetic tasks and pretrained looped language models, reallocation peaks near the largest marginal gain and decays as performance saturates; alternative signals based on entropy or hidden-state change do not consistently exhibit both behaviors. Intervention experiments further support a functional role for this reallocation: freezing attention updates near the reallocation peak causes substantial accuracy loss, whereas removing attention produces a distinct, position-dependent damage profile. This observation yields a training-free stopping rule that stops after the last iteration whose reallocation exceeds 10% of its peak, followed by one margin iteration. Both constants are fixed on synthetic tasks with known required depths and transferred unchanged to pretrained models. The rule matches maximum-depth accuracy on all eight synthetic runs while using roughly half the computation, and transfers across two model families and three scales, saving 25% of FLOPs on Ouro-2.6B and 48% on Huginn-3.5B. On GSM8K and ARC-Challenge, the selected stops satisfy pre-registered equivalence criteria relative to maximum-depth inference. Reallocation profiles are also highly consistent across inputs on the evaluated short-answer tasks, enabling a shared model-level stopping decision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.