Dual-View Anchored Decoding for Mitigating Object Hallucination in MLLMs
Abstract
Multimodal Large Language Models (MLLMs) often suffer from object hallucination, generating objects or attributes that are absent from the visual input. We attribute this failure to the interaction between misleading visual evidence and grounding attenuation: some visual tokens induce uncertain or misleading textual interpretations, while autoregressive decoding progressively amplifies weakly grounded continuations as the textual context grows. Crucially, high epistemic uncertainty does not necessarily imply useless visual information, since a visual token may cover both misleading cues and useful object-specific details. We propose Dual-View Anchored Decoding (DVAD), a training-free decoding strategy that preserves the original full view and constructs an auxiliary anchor view by selectively suppressing visual tokens likely to carry unstable evidence. DVAD jointly decodes the two views with a shared generated prefix and fuses their next-token distributions at each step, thereby suppressing hallucinated continuations without discarding potentially useful information from high-uncertainty tokens. Without retraining, auxiliary models, or stochastic multi-view sampling, DVAD uses a single uncertainty-guided anchor view and can be implemented as a batched dual-stream decode with only marginal wall-clock overhead. Experiments on CHAIR, THRONE, and AMBER across multiple MLLM backbones demonstrate that DVAD consistently reduces object hallucination while preserving caption coverage and quality. Additional evaluations on MMBench and MMStar further show that these gains do not come at the cost of general vision-language understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.