What Matters, What Misleads: Visual Token Attribution and Hallucination Mitigation in LVLMs
Abstract
Large Vision-Language Models (LVLMs) can produce sophisticated multimodal predictions, yet it remains unclear which visual evidence actually drives those predictions. Revealing this evidence is fundamental to understanding LVLM behavior and enabling targeted interventions on the visual information that contributes to hallucinations. However, existing attribution methods based on attention, gradients, or perturbations can overlook computational pathways, provide unreliable estimates, or require costly and context-dependent interventions, motivating a direct approach to measuring the contribution of localized visual-token subsets. In this work, we introduce VIRAH, a training-free framework that retains a local neighborhood of visual tokens at a chosen vision-encoder depth, removes other visual tokens, and propagates the retained representations through the remaining model to measure the evidence supported by selected regions. The resulting output distribution serves as a self-probe of the visual evidence supported by selected regions. We evaluate VIRAH for visual attribution, where the probability of an affirmative answer to a target concept query provides the regional attribution score. We further demonstrate that this self-probing capability can be directly leveraged to develop a training-free hallucination mitigation method, using regional object-identification entropy to identify unreliable visual tokens for removal during captioning. Across various LVLMs (Qwen, InternVL, and LLaVA) and multiple benchmarks, VIRAH consistently outperforms existing attribution methods. Specifically, on Qwen2.5-VL-7B, it improves InsertionDeletion by 46.08% and AUPR by 11.94% over the strongest attribution baseline on MS-COCO. For hallucination mitigation on MS-COCO, this simple training-free intervention achieves performance competitive with or surpassing recent methods, reducing CHAIR by up to 1.90% and CHAIR by up to 2.20% while improving Recall by up to 4.04%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.