Attention-Entropy Decoding with Visual Grounding for Hallucination Mitigation in Multimodal Large Language Models
Abstract
Multimodal large language models (MLLMs) have made significant progress in vision–language reasoning, yet they remain prone to hallucination even when relevant visual information is already present in intermediate layers. Existing mitigation methods either require additional training or address the problem during decoding through visual reinforcement, contrastive prediction, or layer-wise correction. However, accurately identifying which intermediate layer contains useful visual information remains challenging. We hypothesize that the concentration of visual attention provides an effective signal: low-entropy attention indicates concentrated and reliable visual information, whereas high-entropy attention suggests dispersed and unreliable information. Based on this insight, we propose AEDG, a training-free attention-entropy decoding framework that uses visual information to guide the decoding process. At each decoding step, AEDG selects a low-entropy layer as a positive anchor and a high-entropy layer as a negative anchor, and uses them to adjust the final-layer prediction. AEDG requires neither parameter updates nor additional forward computation. Extensive experiments on representative MLLMs show that AEDG effectively reduces hallucination while maintaining strong multimodal capability with only a small inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.