Mitigating Object Hallucination in Large Vision-Language Models via Faith-guided Self-Correcting Decoding
Abstract
Large Vision-Language Models (LVLMs) have demonstrated strong capabilities in open-ended visual understanding, yet remain prone to object hallucinations, generating objects unsupported by the visual input. Existing training-free decoding methods typically correct next-token predictions using auxiliary visual signals, without explicitly assessing the visual support underlying the prediction. We observe that grounded and hallucination-prone predictions respond differently to targeted evidence weakening: grounded predictions tend to lose more predictive support when their attended visual cues are disturbed. Based on this observation, we introduce visual faithfulness, an online measure of visual support derived from intervention-induced prediction sensitivity. We further propose FaiD, a training-free decoding framework that uses visual faithfulness to continuously guide self-correction during generation. Predictions with stronger visual support are preferentially calibrated using positive evidence, while weakly supported ones receive stronger suppression. Extensive experiments across multiple LVLMs and benchmarks show that FaiD consistently reduces object hallucinations while largely preserving grounded content and open-ended response quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.