State-Dependent Visual Support in Vision-Language Hallucination
Abstract
Prior studies have associated vision-language model hallucination with apparently conflicting internal visual signals, including reduced visual dependence and highly attended image tokens. We observe a more specific difference between settings. Hallucinated objects receive less visual support in free-form generation but more when the model makes queried object existence decisions. We show that this difference mainly reflects comparisons of internal representations serving different computational roles. To compare these roles directly, we present StateHall, which applies the same measure of visual support to states corresponding to an object introduced during free-form generation, a queried object, and the queried decision. Hallucinated object representations receive less support around intermediate decoder layers in both free-form generation and explicit querying, while relative over-support appears only at the late decision state. A same-image, same-object decomposition shows that the state-role component constitutes most of the final difference between the generated-object and queried-decision gaps, while the component associated with concept origin is substantially smaller. These patterns are also useful for hallucination detection. For free-form hallucination, the resulting midpoint concept-visual support (MCVS) score matches or improves upon recent training-free baselines and transfers across models and datasets. For queried object existence, compact detectors based on the identified states demonstrate strong transfer across benchmarks and complement output confidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.