acceptodds
Under review as a conference paper at ICLR 2027

Visual Signal Integrity for Detecting Object Hallucinations in LVLMs

Abstract

Large vision-language models (LVLMs) demonstrate remarkable content generation capabilities but are prone to object hallucinations, i.e., generating contextually plausible yet visually ungrounded descriptions. This imposes critical reliability risks in multimodal understanding. Detecting such hallucinations remains challenging because they often arise from breakdowns in visual information preservation that are not captured by existing methods, which inspect only isolated points along the visual-to-language pathway. In this paper, we reveal that grounded object tokens exhibit higher semantic cohesion among attended visual tokens and stronger visual contributions in the residual stream than hallucinated ones. Inspired by these observations, we propose a training-free diagnostic framework, Visual Signal Integrity (VSI), to quantify the integrity of visual grounding during generation. Specifically, VSI introduces two complementary statistics: Visual Semantic Cohesion (VSC), which measures the geometric concentration of attended visual tokens to assess whether coherent visual evidence is formed, and Visual Steering Ratio (VSR), which measures the relative dominance of visual over textual contributions in the residual stream. By combining these two statistics, VSI provides a unified score for detecting breakdowns in visual information flow. Extensive experiments on multiple hallucination benchmarks and diverse LVLMs show that VSI delivers a reliable and interpretable solution for hallucination detection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.