acceptodds
Under review as a conference paper at ICLR 2027

Beyond Visual Attention Magnitude: Visual Sink Reweighting for Mitigating Multimodal Hallucination

Abstract

Multimodal Large Language Models (MLLMs) have demon￾strated strong vision-language capabilities, yet they remain prone to multimodal hallucination, where generated responses are inconsistent with visual content Recent studies have iden￾tified the Visual Attention Sink phenomenon, showing that MLLMs may allocate excessive attention to visually uninfor￾mative or semantically irrelevant visual tokens. While exist￾ing work has analyzed sink tokens and begun to distinguish sinkprone attention heads, the relationship between visual attention magnitude and the structural quality of headlevel visual attention remains insufficiently understood. As a result, highvisualattention heads may be mistakenly regarded as effective visual heads even when they exhibit sinklike patterns, whereas lowvisualattention heads with meaningful visual structures may be overlooked. In this paper, we conduct a structureaware headlevel analysis of Visual Sink and show that visual attention magnitude alone is insufficient for identifying effective visual heads. Motivated by this observation, we propose Visual Sink Reweighting (VSR), a trainingfree method for mitigating multimodal hallucination. VSR constructs a prototypical Visual Sink pattern and introduces SinkScore to measure the structural similarity between each attention head and the prototype. Based on SinkScore, VSR suppresses sink-prone heads while enhancing non-sink heads whose visual attention structures deviate from the Visual Sink prototype during generation. Extensive experiments on multiple benchmarks and MLLMs demonstrate that VSR consis￾tently reduces hallucination and improves visual grounding without additional training or parameter updates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.