Attention Sinks in VLMs Are Borrowed, Not Seen
Abstract
We study how to detect and mitigate visual attention sinks in vision-language models (VLMs). Visual attention sinks are image tokens that absorb a large share of attention over the image regardless of the query or task, and they mislead methods that treat attention maps as a record of where the model “looked.” Prior work has detected sinks through massive activation dimensions and proposed redistributing their attention, but none have addressed what makes a token become a sink. Comparing six candidate detectors on ten VLMs from four families, we find that only the centered cosine similarity between an image token's key and the key of the language model's text-sink detects visual sinks. Thus, a visual sink is an image token whose key aligns with the text-sink's key. Moving a sink's key component along this direction to the median of the image keys sends its attention back to the text-sink (Are Borrowed) while leaving model answers unchanged(Not Seen). Building on this, we propose a training-free, entropy-triggered re-look that first cleans the attention map, then crops the most-attended region of the cleaned map and inserts it into the reasoning trace. Cleaning adds negligible computation to the forward pass, and the re-look improves accuracy with a mean gain of points on fine-detail VQA benchmarks across model families.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.