What Can Be Identified in a Transformer?
Abstract
Mechanistic interpretability names parts of a network (heads, neurons, residual directions), yet many weight settings compute the same function. We ask which of these parts the function determines, and obtain three results. First, a multi-head attention layer is identified exactly: its function determines precisely the measure that assigns to each query–key form the summed value–output maps of the heads using it, already from inputs of length three, so head permutations and per-head gauges are its only symmetries on generic weights. Second, the symmetries of deployed transformers depend on their architecture: rotary embeddings, QK-normalization and attention sinks constrain the gauges, and normalization placement decides whether the residual stream has a privileged basis (none in pre-norm models, only signed permutations when sublayer outputs are normalized). We verify every prediction, with a negative control, on nine checkpoints of seven families up to 32B parameters in float64. Third, these symmetries sort interpretability claims: claims about a checkpoint must be equivariant, claims about its function invariant, and claims about what training produces must replicate across runs. Across independently trained seeds of Pythia and GPT-2, head indices carry no information about IOI or induction attributions, whereas heads matched by symmetry-invariant signatures, and label-invariant statistics, replicate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.