Attending Yet Hallucinating: Inhibitory Evidence Deficits in Vision-language Models
Abstract
Large Vision-Language Models (LVLMs) combine visual understanding with language generation, yet hallucinated responses continue to limit their reliability. Recent studies show that these errors persist even when attention concentrates on relevant visual regions. This raises a question: **why does attending to the right location still lead to an incorrect answer?** We investigate how visual tokens within these regions affect competing answer candidates. Our analysis reveals **Competitive Evidence Imbalance**: (1) Visual tokens within relevant regions provide conflicting evidence, with a substantial share favoring competing answers over the correct one. (2) Tokens favoring the correct answer support it either directly or indirectly by suppressing competing answers. We call the latter **Inhibitory Support Tokens**, whose average count and regional proportion are markedly lower in hallucinated answers despite substantial remaining direct support. Motivated by this finding, we introduce Suppression-prioritized Evidence for Answer Recalibration (SPEAR), a training-free inference method. SPEAR proposes a relevant region from attention, prioritizes inhibitory support tokens while retaining direct support and opposing effects, and aggregates their signed responses to calibrate candidate probabilities under a bounded distributional change. Experiments show improved object hallucination mitigation, compositional discrimination, open-ended answer quality, and aggregate visual perception. These results highlight the value of using visual evidence to suppress competing answers as well as support plausible ones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.