Adaptive Visual Evidence-Guided Decoding for Hallucination Reduction in LVLMs
Abstract
Large Vision-Language Models (LVLMs) are prone to hallucinations, producing outputs that are linguistically plausible but not grounded in visual evidence. Existing mitigation approaches typically rely on multi-pass interventions, attention reallocation, or look-twice visual re-injection mechanisms. Although effective, these methods either incur significant computational overhead or provide only partial mitigation, limiting their reliability in practical deployment. In this work, we introduce Adaptive Visual Evidence-Guided Decoding (AVED), a unified single-forward-pass framework that explicitly incorporates visual evidence into the decoding process. Instead of treating visual grounding as an external correction step, AVED integrates visual evidence estimation and decoding within a single forward pass, eliminating the need for multiple inference stages or external visual re-injection. Specifically, our method consists of two complementary components: Visual Attention Calibration (VAC), which refines cross-modal attention to capture vision-language correspondence, and Layer-wise Logits Reweighting (LLR), which adaptively aggregates intermediate decoding predictions according to their estimated visual evidence reliability. The two components are tightly connected through cross-attention signals, where calibrated attention provides visual evidence estimates that guide evidence-aware logits reweighting. Extensive experiments on hallucination benchmarks demonstrate that our approach consistently suppresses hallucinations while improving generalization under distribution shifts, establishing a simple yet effective paradigm for reliable LVLM inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.