acceptodds
Under review as a conference paper at ICLR 2027

Attention Sink Functionality Is Query-Conditional

Abstract

Attention sinks are usually treated as a single phenomenon, yet the literature assigns them opposite roles: near-uninformative rest states in some accounts, functional carriers of information in others. We argue that this conflict is partly a coordinate problem: attention mass, value norm, and Euclidean value geometry do not directly measure what a sink contribution does to the output. We introduce \method brightness, the variance a residual direction induces in the next-token distribution through the unembedding, and use it to audit first-token sink contributions in decoder-only language models. Our central finding is that sink output functionality is query-conditional. The same first-token value direction, in the same head and text, can strongly perturb one query's next-token distribution while remaining quiet for another. In matched-dose event pairs, attention-mass and residual-norm log-ratios stay within , yet readout brightness can differ by more than two orders of magnitude; brightness stratifies which suppressions perturb the local output, with model-dependent strength. The strongest direction-specific evidence comes from Mistral-7B and a seed-split-selected Qwen3-8B layer-28 site, where a strict fp32 rerun gives a local-KL gap and exceeds same-norm random-direction actual interventions by . Llama-3.1-8B is positive but weaker at the pre-specified layer-20 site; a fixed-candidate audit over already-audited layers identifies a stronger layer-31 site. Gemma2-9B further shows that the local effect survives a soft-capped-readout stress test. Finally, a normalization-aware small-perturbation audit separates the Fisher diagnostic from the finite-dose causal effect: the exact direct-readout counterfactual converges to the RMSNorm-corrected Fisher prediction as , while finite-dose interventions measure propagated network effects. These results suggest that single-role accounts of attention sinks partly reflect coordinate choice: output-quiet sink components are not globally irrelevant, but their output function is visible only in an output-readout coordinate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.