SSA: Scaled Sink Attention for Large Language Models
Abstract
As large language models (LLMs) attract growing interest, the attention mechanism at their core has become one of their most critical components. Yet Softmax attention still exhibits attention sink: a large share of the probability mass is placed on a few uninformative tokens, usually the first one. Recent remedies relax the sum-to-one constraint of Softmax with a learnable denominator bias (Sink Attention) or a sigmoid output gate (Gated Attention) and report much lower first-token scores. We point out, however, that any such per-query rescaling, including ours, leaves the normalized attention distribution unchanged for fixed logits; whether a method removes the sink is decided by what the trained model learns and must be measured after normalization. Under this measure, existing remedies still leave a substantial share of the attention on the first token and, in addition, suppress attention on informative tokens. We propose (SSA), which combines two components: (i) an input-dependent virtual sink, a per-head logit projected from the hidden state that competes inside the Softmax and is then discarded, so that the drained mass is decided token by token against the attention evidence; and (ii) a fixed gain on the surviving attention, parameterized by a reference surviving mass , which restores the output scale of heavily draining heads and sharpens contextual attention elsewhere. SSA substantially lowers the normalized first-token share and redistributes attention to informative context tokens. SSA is also practical: it adds only one projection per layer, and we have integrated it into FlashAttention-3 as a ready-to-use fused kernel, so that it inherits the efficiency of state-of-the-art attention kernels and incurs an end-to-end training slowdown of less than 1%. In extensive experiments on a from-scratch pre-trained Mixture-of-Experts model, SSA improves the fourteen-benchmark average by 2.74 points over Softmax attention, reaches the baseline's final pre-training loss with about fewer tokens, and never triggers the QK-clip safeguard that all baselines trigger hundreds to thousands of times.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.