acceptodds
Under review as a conference paper at ICLR 2027

How Weight Tying Elicits Iterative Computation in Looped Transformers

Abstract

Looped transformers apply the same block of layers many times to each input, and have been shown to learn iterative solutions that progress roughly one step per iteration (a staircase). It remains unclear, however, how exactly weight tying leads to such solutions. We investigate this using directed permutation traversal, a task in which a model must report the node reached after hops from a starting node. Its precisely defined intermediate states let us track the model's progress iteration by iteration and layer by layer using linear probes and causal interventions. We find that the underlying mechanism is gradient pooling, the summing of gradients across repeated uses of the shared block. We test this by giving each iteration its own copy of the block and varying how much of each copy's update comes from the pooled gradient. With more pooling, the task is learned more quickly and reliably, and the solution moves from the layout of untied models, which expose intermediate states only near the output, to the staircase. Conversely, training a looped model on only one iteration's gradient, however it is scaled or sampled, prevents the solution from forming. Pooling thus favours a reusable hop operation which, once formed, is composed across iterations, explaining both the rapid acquisition of multi-hop capability during training and the staircase at inference. Together, these findings provide a mechanistic account of how weight tying elicits iterative computation, informing design choices as looped models scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.