Exploring and Mitigating Object Hallucination with Non-Autoregressive Vision-Language Models
Abstract
Vision-Language Models (VLMs) support a growing range of multimodal applications, yet their generated descriptions remain prone to object hallucination. A conservative caption can avoid unsupported mentions by saying less, but the resulting loss of coverage limits usefulness; this tension appears clearly across representative generation paradigms. Autoregressive VLMs such as LLaVA-1.5 describe more of the scene but hallucinate more objects, whereas non-autoregressive masked-diffusion VLMs such as LLaDA-V achieve much lower hallucination but leave more visible content unmentioned. We trace this contrast to how the two generation architectures resolve uncertainty: autoregressive decoding makes each emitted token an immediate, irreversible part of the context, whereas masked diffusion can revise low-confidence predictions before fixation. We formalize this distinction as *selective object commitment*, separating broad object discovery from the final decision to assert each candidate. Specifically, we decompose final commitment into proposal admission, diffusion realization, and omission recovery, and derive the stagewise selectivity needed to expand coverage without proportionally increasing hallucination risk. Guided by this formulation, we introduce COVER (Cross-view Object Verification and Evidence-guided Recovery), a training-free method that combines autoregressive and diffusion proposal views, builds an object plan from cross-view agreement and object-level real/null visual gain, applies the same contrast through Visual Residual Guided diffusion decoding (VRG) to guide token selection and fixation, and restores verified omissions through bounded recovery. On MSCOCO evaluations, COVER reduces sentence-level CHAIR by 28.1% while retaining 98.7% of the detailed proposer's object Recall. COVER establishes proposal-commitment separation as a practical cross-architecture principle for grounded multimodal generation. *The source code for our method will be made available upon publication.*
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.