AMEND: Adaptive Modulation via Enhanced Navigation Decoding for Mitigating Hallucinations in LVLMs
Abstract
Large Vision-Language Models (LVLMs) have demonstrated extraordinary capabilities in cross-modal tasks. However, they frequently suffer from hallucinations, generating descriptions that are inconsistent with the visual inputs. Existing training-free attention intervention methods lack dynamic adjustment, resulting in poor visual grounding or instruction neglect, while indiscriminately enhancing all visual tokens amplifies background noise. To address this, we propose AMEND, a novel training-free approach that mitigates hallucinations by dynamically recalibrating visual attention. Specifically, this adjustment is performed from two aspects: (1) Uncertainty-Aware Adaptive Scaling, which leverages the entropy of output token logits to estimate generation uncertainty and adaptively adjusts the visual attention enhancement coefficient; and (2) Expert-Layer Attention-Driven Token Extraction, which analyzes attention weights across vision-dominant expert layers to selectively identify the most critical and question-relevant visual tokens, rather than indiscriminately amplifying all image tokens. Finally, using the dynamically enhancement coefficient, we amplify attention to the selected key visual tokens while suppressing attention to irrelevant visual tokens and system prompts. Extensive experiments on standard benchmarks demonstrate that AMEND significantly alleviates hallucinations and outperforms existing baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.