When Is an Attention Head a Computational Unit? Shared Circuits Despite Orthogonal Representations
Abstract
Mechanistic analyses of transformers often proceed head-by-head, looking for attention maps that correspond to computational roles. But when should an attention head be interpreted as a computational unit? In principle, a single computational function may be distributed across heads, and a single head may participate in several computations, with invariants appearing only after OV routing and residual-stream summation. In language models, where the underlying features and computations are unknown, an apparently uninterpretable head may be polysemantic, part of a distributed circuit, or a sign that the analyst chose the wrong decomposition. We therefore study transformers trained on factored sequence processes, a controlled setting where the target computation is analytically specified: multiple latent factors, represented in orthogonal residual-stream subspaces, receiving independent predictive updates. This lets us ask whether the attention circuits implementing these updates are themselves factorized. We find that they need not be. Per-factor updates constrain only the aggregate contribution across heads, leaving substantial freedom in how computation is distributed. Heads may specialize to individual factors, compose with other heads to implement one factor, or contribute polysemantically across factor boundaries. The regime that emerges depends on the generator's spectrum, the number of available heads, and training dynamics. We introduce effective subspace attention, a scalar quantity that combines attention patterns with OV circuits to measure the amount that heads collectively route from a source position into a given factor subspace. In our controlled setting, this aggregate view recovers the theoretically predicted computation even when per-head maps are illegible. Because only the aggregate is constrained, computation may be best understood at the level of task-relevant subspaces rather than individual heads. We then ask whether effective attention can support circuit discovery in real models and find that its natural application is closely related to existing circuit analysis techniques. When applied to GPT-2, it recovers known and novel circuits on two contrastive next-token tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.