acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Contrastive Decoding: Evidence Guided Branch Assignment for Multimodal Hallucination Mitigation

Abstract

Multimodal large language models (MLLMs) integrate visual perception with the language understanding and generation capabilities of large language models, enabling a wide range of multimodal tasks. However, object hallucination remains a critical challenge that undermines the reliability of MLLMs. Existing contrastive decoding methods mitigate hallucinations by constructing positive and negative branches through visual perturbation, token selection, or attention modulation. However, the roles of these branches are often determined by predefined operations rather than by explicitly verifying whether their contents are supported by the input image. To address this limitation, we propose Visual Evidence guided Role Assignment (VERA), an inference time hallucination mitigation framework that verifies intermediate visual hypotheses before constructing contrastive branches. VERA assigns visually supported components to the positive branch, while uncertain hypotheses are reverified before their unresolved residuals are used for negative suppression. To prevent residual information from being inadvertently converted into positive evidence during contrastive subtraction, VERA further constrains the residual branch to provide suppression only. Extensive experiments on CHAIR, POPE, and five general visual question answering benchmarks demonstrate that VERA effectively reduces object hallucinations while preserving the general multimodal understanding capabilities of MLLMs across LLaVA-1.5, LLaVA-NeXT, and Qwen2.5-VL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.