acceptodds
Under review as a conference paper at ICLR 2027

How Training Loop Counts Shape Computation in Looped Transformers

Abstract

A looped transformer applies the same block repeatedly before producing an answer. We ask how the number of loops used during training changes the computation learned by that block. We study sequence tasks with a known correct output at each position. Reading these outputs after each loop reveals a computation frontier that advances from the start of the sequence. Its average advance per loop is frontier speed. For cumulative products of permutations of four objects, models trained with fewer loops show faster measurable frontiers. This relation holds when models share initial weights and training examples and receive the same number of optimizer updates. It also appears when accuracy determines advancement through a length curriculum. We then continue training two copies of each source model. Reducing the loop count increases frontier speed relative to keeping it fixed. Replacing intermediate states disrupts later outputs. Additional loops can extend correct computation to the last position and then reduce the fraction of fully correct answers. These findings show how training shapes progress inside a shared block and how that progress determines when its output is useful.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.