A Unifying View of Attention Sinks: From Mechanisms to Architectural Interventions
Abstract
When attention concentrates on a single token, a *sink*, what is the model actually computing? Attention sinks are ubiquitous in softmax transformers, yet this shared visual signature can hide fundamentally different algorithms. We show that visually similar sink patterns can reflect two distinct mechanisms: *(i)* adaptive *NOP*, where a head suppresses its update by routing to a null token, and *(ii)* broadcast, where a sink aggregates and redistributes global information. Each mechanism leaves distinct traces (*NOP* sinks exhibit negligible value norms; broadcast sinks induce low-rank outputs), which we formalize on synthetic tasks and use to derive practical diagnostics. Applied to pretrained vision transformers, these diagnostics reveal that both mechanisms exist at scale: sinks transition from `CLS` in early layers to patches in deeper layers and concentrate in specialized heads. Causal interventions further connect these signatures to near-null suppression and shared residual contributions. We then use architectural interventions to show how these computations can be reorganized: gating eliminates detected *NOP*-like sinks but increases broadcast-like sinks, registers relocate rather than remove sink computation, and our position-free global pathway provides an explicit route for shared communication that reduces the broadcast-like sinks induced by gating. On dense probes, combining gating with the global pathway gives the strongest results among the tested variants, despite retaining some broadcast-like sinks. Overall, we find that the same attention pattern can reflect two very different computations, and that effective intervention depends not only on identifying the computation, but also on providing architectural alternatives through which the model can reorganize it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.