From Abnormal Visual Tokens to Faithful Decoding: Rebalancing Visual Evidence for Hallucination Mitigation in MLLMs
Abstract
Multimodal Large Language Models (MLLMs) remain vulnerable to hallucinations despite their strong visual understanding and generation capabilities. Existing hallucination mitigation methods mainly focus on language-side biases, while abnormal visual behaviors and corresponding theoretical implications for faithful decoding remain largely underexplored. Through extensive and statistically significant investigations into the causes of hallucination, we reveal that unusually high-magnitude visual tokens exhibit persistent off-target responses and increasingly interfere with grounded object prediction in deeper layers, accompanied by a progressive weakening of visual evidence during decoding. Based on these observations, we propose a Visual Evidence Rebalancing and Grounding Enhancement method called VERGE, a training-free framework that combines evidence-guided channel correction, uncertainty-triggered visual retracing, and prior-calibrated contrastive decoding. Furthermore, we provide theoretical analyses that characterize the improvement of faithful-token preference and the benefit of recovering target-relevant visual information. Extensive comparative experiments demonstrate that VERGE effectively suppresses hallucinations while maintaining strong generation capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.