acceptodds
Under review as a conference paper at ICLR 2027

Attention Sinks in Transformer-based Language Models: Emergence Mechanisms and Benign Effects

Abstract

Attention sinks, in which large language models (LLMs) disproportionately attend to specific tokens such as the initial token or weak semantic tokens like punctuation and function words, have become an important focus in understanding Transformer mechanisms. This phenomenon has been shown to affect model performance and efficiency, motivating methods that either mitigate or leverage it to improve architectures. However, its theoretical underpinnings remain insufficiently understood, particularly how these patterns emerge spontaneously during training. In this work, we address this gap by theoretically analyzing a two-block Transformer trained on synthetic text datasets. We show that causal masking causes parameters associated with the first token to accumulate disproportionate gradients, driving the formation of a first-token attention sink, while high-frequency semantic tokens receive significantly larger gradients than other tokens and consequently absorb most of the attention. We further show that these attention sinks are compatible with the Transformer's predictive capacity, even when attention is heavily concentrated on a few tokens. Overall, this work provides a training dynamics explanation for the emergence of attention sinks, offering insights that inform the design of new algorithms and architectures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.