Identifying the Functional Distance Between Attention Heads
Abstract
Attention heads in transformers are routinely compared through the similarity of their attention rows. Most heads place their largest attention weight on a sink token, and analysts must choose whether to keep or remove the sink before comparing, without knowing what either choice measures about a head's effect on the model. We prove that attention rows do not determine how differently two heads act, while the value vectors do. We show this by writing the difference between two head outputs as the value matrix applied to the difference of their attention rows: the rows leave the distance free within an interval, and the Gram matrix of the values fixes it exactly. We then introduce the recipient-slot distance, the output change when one head's output passes through the output projection of another, and prove that this substitution changes the layer output less than removing the head exactly when the distance is below the head's output norm. On GPT-2, Pythia, and Llama, comparisons based only on attention rows agree with the value-aware comparison about as often as chance, and about nine in ten heads change their functional nearest neighbour. Head outputs within a layer are nearly orthogonal, so their distance is set mostly by their norms, which attention rows do not contain. Selecting heads for removal with the value-aware distance costs less perplexity than selection by attention rows in eight of nine settings, and on Llama, where heads share values, substitutions chosen by the recipient-slot distance cost less than every removal rule. We further split every standard dissimilarity exactly into a sink part and a content part. This split shows when keeping or removing the sink reverses conclusions. Under natural axioms, no dissimilarity can be exactly separable and also continuous at the boundary or monotone under processing. Practitioners can compare heads within a layer through their projected values and report sink mass and content distance separately.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.