When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
Abstract
Multimodal large language models (MLLMs) have become a key interface for visual reasoning and grounded question answering, yet they remain vulnerable to visual hallucinations, where generated responses contradict image content or mention nonexistent objects. A central challenge is that hallucination can coexist with substantial visual attention: the model may assign considerable attention mass to image tokens while internally drifting toward an incorrect answer. In this paper, we find that the high-frequency structure of visual attention, measured by layer-wise Laplacian energy, helps select intermediate layers with contrasting preferences for hallucinated and correct answers. Building on this finding, we propose LaSCD (Laplacian-Spectral Contrastive Decoding), a training-free decoding strategy that selects informative layers via Laplacian energy and remaps next-token logits in closed form. Experiments on hallucination and general multimodal benchmarks show that LaSCD consistently reduces hallucination while preserving general capabilities, highlighting its potential as a faithful decoding paradigm.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.