acceptodds
Under review as a conference paper at ICLR 2027

CAVCD: Candidate-Aware Visual Counterfactual Decoding for Mitigating Object Hallucinations in Large Vision-Language Models

Abstract

Large vision-language models (LVLMs) suffer from object hallucination, where generated responses mention objects absent from the input image. Contrastive decoding alleviates this problem by contrasting original and perturbed predictions, but offers limited control over which candidate tokens to correct and how strongly, potentially disrupting otherwise correct predictions. To address these limitations, we propose Candidate-Aware Visual Counterfactual Decoding (CAVCD), a training-free method built on the idea that intervention-induced prediction changes should be assessed for individual candidates before guiding token selection. CAVCD first adapts visual interventions to the generation process, restricting visual access at the initial step and applying softer attention attenuation during subsequent generation. During continuation, it evaluates candidate responses across decoder layers and combines selected responses into a target distribution that guides logit correction. An adaptive gate then regulates correction strength, limiting excessive updates that could override image-supported token choices. Together, these designs make correction more selective and adaptive, helping suppress hallucinations while preserving plausible predictions. Extensive experiments demonstrate that CAVCD effectively mitigates object hallucination across multiple evaluation metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.