Attention Sinks Beyond the First Position: Heterogeneous Pathways
Abstract
Attention sinks, in which a small subset of tokens in a sentence receives disproportionately large attention mass, are commonly associated with the first position. However, tokens at non-initial positions can also become attention sinks, yet the circuit mechanisms underlying these secondary sinks remain poorly understood. In this work, we investigate their formation in the first Transformer block using complementary activation-patching experiments. Across the evaluated models and token types, secondary sinks arise through different contributions from word embeddings and self-attention outputs. When the "bos" token appears at a non-initial position, its sink behavior is driven predominantly by its embedding in some models and by a combination of its embedding and attention output in others. For the other tested secondary sink tokens, their sink behavior relies mainly on the strong self-sinking within a small subset of critical heads in the first-block attention output. Further experiments reveal a corresponding difference in sensitivity to contextual mixing: embedding-dominated sinks persist under contextual mixing, even in the presence of competing secondary sink tokens, whereas attention-output-dominated sinks can disappear when competing tokens occur earlier in the sequence. These findings suggest multiple first-block pathways to secondary sink formation and show how these pathways shape the sink behavior under different contexts. More broadly, our analysis provides a circuit-level framework for understanding and distinguishing attention-sink mechanisms in large language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.