Sink or Swim: Diagnosing and Resolving the Relational Invariance Bottleneck in Causal Attention
Abstract
Decoder-only Transformers consistently concentrate attention on the first token, a phenomenon known as attention sink. We trace its origin to a structural constraint we call the relational invariance bottleneck: causal attention has no neutral gear— row-stochasticity forces every head to move mass somewhere, while causality makes the first position the only universal sink. An attention matrix is a universal relational invariance if it preserves every input up to a token-shared shift; under this definition, we establish the exact intersection: Collapse onto token one is therefore not an incidental artifact but the only available invariance target. To quantify this mechanism, we introduce a relational energy measuring how much token-level variation a head’s output carries, and prove that it is bounded by the head’s distance to the invariance target: This separates cause from symptom: first-token concentration matters only insofar as it signals proximity to the invariance target that suppresses relational output. Motivated by this diagnosis, we propose Adaptive Diagonal Target Editing (ADTE), a lightweight per-head module that cancels the self-token component of each attention row and rescales the remaining cross-token signal, giving heads an alternative relational invariance route that no longer requires concentrating mass on the first token. Across two architectures, ADTE consistently lowers pretraining perplexity compared with all baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.