acceptodds
Under review as a conference paper at ICLR 2027

The Anchoring Effect of Attention Sinks

Abstract

Attention Sinks (AS), tokens that accumulate disproportionately high attention weights, are commonly regarded as detrimental artifacts in Large Vision-Language Models (LVLMs), often associated with hallucination and attention redundancy. In this work, we show that this view is incomplete: AS also serve as structural anchors in LVLMs. Specifically, AS persist across layers and attention heads, while targeted perturbations induce disproportionate changes in attention organization and representation geometry. Further masking experiments reveal their dual role: Masking sink attention can mitigate hallucinations, but this benefit may come at the cost of degraded visual reasoning and broader model capabilities. Motivated by these findings, we propose Gated Sink Attention Redistribution (G-SAR), a simple input-adaptive gating mechanism that learns how much attention to release from sink tokens and redistributes the released mass toward visual non-sink tokens. Across three LVLMs and diverse general visual understanding, hallucination, and visual reasoning benchmarks, G-SAR consistently improves overall performance while reducing hallucination. Our results suggest that AS should be regulated rather than eliminated, highlighting adaptive sink control as a more principled strategy to balance visual grounding and reasoning in LVLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.