Identity Supervision for Readable Intermediate States
Abstract
How do looped transformers organize intermediate states that new downstream computations can use across source views? We investigate this question through controlled studies of modular addition and two-hop relational composition, comparing weight-tied and untied transformers at matched effective depth. Our analysis combines temporary supervision interventions, state diagnostics across checkpoints, and tests of new readers and retained downstream computation. These tests distinguish identity information readable from a state from its usefulness to a specified attribute reader, and examine the roles of supervision location, gradient exposure and source variation. The tested identities are known, and the later reader generalizes their observed labels across source views. In relational composition, direct identity supervision yields gains of 5.2–7.6 percentage points across seven settings under a common bounded gradient rule, while several comparisons with fixed calibration remain unresolved. In modular addition, near-perfect affine identity reading coexists with substantially weaker random-attribute reading. Depth-by-strength comparisons qualify a pure supervision-distance explanation, and matching identity grouping reveals different costs for source-factor reading. Together, these results characterize concrete dependencies and limits of intermediate-state use rather than a universal benefit of weight sharing. The study connects source-view generalization to analysis of internal states while distinguishing reader-specific accessibility from stronger claims about computational mechanisms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.