Same Head, Different Calls? Auditing Functional Role Claims in Looped Transformers
Abstract
Looped Transformers apply one block across calls within a forward pass. Its weights are shared, yet its recurrent states evolve: a changing intervention effect can therefore invite a claim that a head changes function. What does such a profile establish? We specify the recipient states, donor rule, readout horizon, and response construction, pairing each claim with a diagnostic comparison. In pretrained Ouro-2.6B, a selected head shifts the answer margin at a trained early exit by 0.72 logits, yet the effect vanishes after the next call. Across six behavior-selected pointer models, target-unrelated donors reproduce 46% of the matched patch contrast and show progress-aligned structure while promoting their own states. Matched donors additionally promote the intended clean state: contrast improvement and target promotion are distinct. In separate vector analyses, fitted phase gains depend on response coordinates, and the role-slot gain weakens when the clock represents the task boundary. Sparse within-clock phase variation and head-weighting sensitivity leave phase information beyond time unresolved. In a cued positive control, scalar checks retain a known terminal effect. The evidence supports transient influence and donor-conditioned target promotion, but does not identify functional role switching.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.