acceptodds
Under review as a conference paper at ICLR 2027

Improving Test-Time Scaling with Adaptive Looped Transformers

Abstract

Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy–compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. Our token-level analysis shows that fixed-depth looping spends extra iterations on every token, although many tokens do not benefit from them. We therefore propose TaH2, which enables the model to focus extra iterations on the tokens that benefit from looping. It jointly post-trains the backbone and an iteration decider through lookahead depth supervision, with online labels indicating whether further iteration improves the prediction. A learned updater reinjects token embeddings between iterations, while the decider's stopping probabilities weight predictions across the executed depths. TaH2 improves both the efficiency and attainable accuracy of test-time scaling. On challenging AIME benchmarks, TaH2 improves the accuracy–compute slope by 53% (2.74 vs. 1.79) over the non-looped baseline, exceeding the baseline's peak accuracy by about 3.4 points at matched compute. As the maximum iteration depth increases, existing looped models largely plateau, while TaH2's gain over the non-looped baseline grows from +2.8 points at depth 2 to +3.9 points at depth 8.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.