acceptodds
Under review as a conference paper at ICLR 2027

Reliable Memory Injection into Core Visual Attention Heads for Object Hallucination Mitigation in Large Vision-Language Models

Abstract

Large vision-language models (LVLMs) often hallucinate objects that are absent from the input image during open-ended generation. A primary cause is the progressive weakening of visual evidence during autoregressive decoding, which allows language priors to dominate. Existing methods typically enhance visual information at the model or layer level while largely overlooking functional differences among attention heads. To address this limitation, we first analyze visual perception in LVLMs from the perspective of multihead attention and identify Core Visual Attention Heads (CVAHs) through statistics collected across multiple samples. We then introduce Core Visual Attention Head Memory Injection(CAMI), a training-free method. CAMI uses relative visual-attention decay as a state-dependent intervention scheduler, not as a classifier of hallucinated tokens. When triggered, it injects historical visual pattern memory to correct the spatial attention of CVAHs. Experiments on MSCOCO-CHAIR and AMBER-G show that CAMI consistently mitigates object hallucination and outperforms existing state-of-the-art baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.