When Visual Truth Loses Control: Diagnosing and Preserving Grounded Predictions in Multimodal Models
Abstract
Reliable multimodal generation requires visual evidence to remain influential throughout decoding, yet existing studies mainly attribute hallucination to insufficient visual perception or excessive language priors, leaving its internal evolution underexplored. We uncover a fundamental failure mode, termed Grounded Control Reversal (GCR), where a visually grounded preference emerges in intermediate representations but is subsequently weakened or overturned before the final answer is produced. Through layer-wise preference tracing and same-image internal interventions on Qwen2.5-VL-7B across HallusionBench and POPE-Adversarial, we show that reversals arise from the combined effects of attenuated visual pathways, competing textual pathways, and feed-forward transformations. Based on this insight, we introduce ReCAP, a training-free decoding framework that identifies a persistent visually grounded intermediate state and restores its lost predictive influence through a minimum-norm projection in hidden space. The intervention selectively recovers the grounded preference while preserving unrelated residual components. Our findings provide a causal perspective on multimodal hallucination, showing that failures can arise not from the absence of visual understanding, but from the subsequent loss of control over already acquired visual evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.