TAVCD: TARGET-AWARE VISUAL CONTRASTIVE DECODING FOR HALLUCINATION MITIGATION IN VISION- LANGUAGE MODELS
Abstract
Visual Contrastive Decoding (VCD) mitigates object hallucination in Large Vision-Language Models (LVLMs) by contrasting positive-branch logits (original image) against a negative branch (degraded image), weighted by a fixed scalar α. We expose a fundamental limitation of this paradigm, formalised as the Contrastive Correctability Condition (C³): no fixed positive α can correct a hallucinated token h whenever h simultaneously dominates the ground-truth token g in both the baseline logit distribution and the visual contrastive signal D(v)=ℓ⁺(v)−ℓ⁻(v). We further prove a universal insufficiency result: for any fixed α₀∈ℝ there exist realizable decoding configurations under which α₀ either fails to correct or actively amplifies hallucination. Empirically, across several backbones on MSCOCO, C³ is violated at approximately 26% of object-generating steps, confirming that fixed-α VCD is counterproductive at roughly one in four critical moments. Building on this theory, we propose TAVCD, a lightweight per-step adaptive contrastive decoder. At each decoding step TAVCD builds a compact 10-dimensional feature vector from observable logit statistics and routes it through a shared-trunk gate+magnitude MLP (≈12K parameters). The gate decides whether contrastive intervention is warranted; the magnitude head selects the signed strength from a fine-grained 40-point α grid. TAVCD is trained offline on logged decoding traces with ground-truth onset supervision and adds negligible wall-clock overhead at inference. On MSCOCO and PASCAL VOC captioning across five diverse LVLMs, TAVCD reduces hallucination on the majority of backbones, achieving up to 24.0 pp CHAIRᵢ improvement over standard decoding and up to 5.8 pp CHAIRₛ improvement over fixed-α VCD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.