acceptodds
Under review as a conference paper at ICLR 2027

Reverse Engineering Attention Sinks in LLMs

Abstract

In autoregressive large language models (LLMs), the beginning-of-sequence (BOS) token tends to receive disproportionate amounts of attention. While past research has theorized about why these attention sinks form, we instead focus on characterizing what attention sinks are and how they're implemented. We posit that attention sinks are used to implement linear threshold-based gates to filter attention logits, and from this, a Heisenberg uncertainty principle follows: an attention head can either observe an attention sink, or observe its semantic meaning and position, but not both simultaneously. Inspired by this principle, we find that LLMs trained to accept sequences without a BOS token, which have been shown to form an attention sink at their first token, use Layer 0 to remove semantic and positional information from the first token, turning it into a replacement BOS so that it may function as the attention sink. We show how malicious prompts can fool these LLMs into acting as if each of multiple tokens are the unique attention sink. Finally, motivated by our findings, we propose minimal architectural modifications to alleviate these issues, resulting in LLMs that are comparable-or-better in performance, no longer form first-token sinks, and are more robust to repeated-token attacks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.