acceptodds
Under review as a conference paper at ICLR 2027

WHAT DO VISUAL-HEAD DISCOVERY SCORES ACTU- ALLY FIND? ATTENTION SOURCE, NECESSITY, AND RANDOM-CONTROL PITFALLS IN VISION-LANGUAGE MODELS

Abstract

Attention heads that look at the image region a question refers to are increasingly used to steer, ground and interpret vision-language models (VLMs) and to reduce their hallucinations. They are mostly found by ranking heads with an attentionpattern score and validated by steering or masking them. We ask how the choices inside such a discovery score shape the heads it returns, and which of the two validations identifies them. On Qwen3-VL-8B-Instruct we vary the query type, the discovery corpus and the token positions whose attention is read (the attention source) in a full 2 × 3 × 2 design, and test all 12 head-sets for steerability (CERS) and necessity (NCS) against layer-matched random head-sets. (1) The attention source and the query type reshape the selected heads far more than the corpus: changing the source keeps a median of 49 of the top-100 heads, the query 55 and the corpus 84, while two halves of the same data share 96–99. (2) Necessity follows the attention source: head-sets discovered from query-span attention are more necessary than last-token ones (difference 0.083, p=6×10−10), and the two referring + query-span sets lie 17–19 standard deviations above 250 layer-matched random sets (Holm-adjusted p=0.048). The referring + query-span effect replicates across several recent VLMs. (3) Steerability is not specific to the discovered heads: no discovered head-set steers better than its layer-matched random sets (rank p ≥ 0.14), and random sets drawn from the same or from any layer span the discovered sets’ scores. Necessity therefore separates the discovery settings, while steering at this strength does not. We turn these findings into a protocol for reporting visual-head discoveries.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.