acceptodds
Under review as a conference paper at ICLR 2027

When Visual Truth Loses Control: Diagnosing and Preserving Grounded Predictions in Multimodal Models

Abstract

Reliable multimodal generation requires visual evidence to remain influential throughout decoding, yet existing studies mainly attribute hallucination to insufficient visual perception or excessive language priors, leaving its internal evolution underexplored. We uncover a fundamental failure mode, termed Grounded Control Reversal (GCR), where a visually grounded preference emerges in intermediate representations but is subsequently weakened or overturned before the final answer is produced. Through layer-wise preference tracing and same-image internal interventions on Qwen2.5-VL-7B across HallusionBench and POPE-Adversarial, we show that reversals arise from the combined effects of attenuated visual pathways, competing textual pathways, and feed-forward transformations. Based on this insight, we introduce ReCAP, a training-free decoding framework that identifies a persistent visually grounded intermediate state and restores its lost predictive influence through a minimum-norm projection in hidden space. The intervention selectively recovers the grounded preference while preserving unrelated residual components. Our findings provide a causal perspective on multimodal hallucination, showing that failures can arise not from the absence of visual understanding, but from the subsequent loss of control over already acquired visual evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.