acceptodds
Under review as a conference paper at ICLR 2027

Beyond Visual Attention: Competition-Informed Visual Evidence Refinement for Hallucination Mitigation in LVLMs

Abstract

Hallucination remains a critical challenge for Large Vision-Language Models (LVLMs), where generated responses may be fluent yet inconsistent with the actual visual content. Existing studies commonly attribute hallucination to insufficient visual contribution and mitigate it by strengthening visual signals during decoding. However, hallucination may still occur even when the model attends strongly to the image. We find that hallucinated predictions exhibit greater overlap between the visual-evidence distributions associated with competing token candidates, suggesting that the attended visual information may not sufficiently distinguish these candidates. We term this phenomenon candidate-level visual-evidence ambiguity and find that it is associated with a higher rate of hallucination. Based on these findings, we propose Competition-Informed Visual Evidence Refinement (CIVER), a training-free decoding framework that disambiguates competing token predictions using four complementary components: a sample-specific visual prior, supporting evidence for the provisional leading candidate, competing evidence favoring its strongest alternative, and cross-step evidence memory. CIVER combines these components into a signed evidence bias and selectively applies it to subsequent visual attention according to the ambiguity of the candidate pair. Extensive experiments on POPE, CHAIR, MME, and AMBER using InstructBLIP, Shikra, LLaVA-1.5, and Qwen-VL demonstrate that CIVER consistently reduces hallucinations and improves visual faithfulness across diverse LVLM architectures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.