Seeing Before Believing: Mitigating Hallucinations in Multimodal Reasoning via Visually-Anchored Latent Rectification
Abstract
Although long-chain reasoning aids complex multimodal tasks, the extended thinking trajectory can drift from image evidence and produce hallucinated answers. Recent work attempts to rectify reasoning by optimizing latent states guided by reference-free predictive confidence. However, predictive confidence becomes misleading when greater certainty reflects linguistic predictability rather than visual faithfulness. When reasoning detaches from the image, self-reinforcing language priors can concentrate probability on predictable continuations, producing a bungee-like rebound in confidence and a false sense of certainty without visual support. Optimizing against confidence alone thus promotes visually ungrounded shortcuts, amplifying such high-confidence errors. To address this, we introduce Visually-Anchored Latent Rectification (VALR), a training-free method that uses visual conditioning to guide confidence-based latent refinement. VALR dynamically retrieves and fuses aligned image evidence into candidate states, shaping the reward landscape around visually induced confidence gains to guide exploration and select the state used for answer generation. Across three models, VALR outperforms all compared inference-time methods in hallucination mitigation, while its benefits extend to broader multimodal understanding and mathematical reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.