acceptodds
Under review as a conference paper at ICLR 2027

Masked-State Evidence for Correcting Object Hallucinations in Diffusion Vision-Language Models

Abstract

Object hallucinations remain common in diffusion vision–language models (dVLMs). Unlike autoregressive VLMs, which expose a causal prefix that can be rolled back, dVLMs resolve a masked response canvas in parallel. Consequently, commitments need not follow a unique order, and completed bidirectional context can reconstruct an unsupported semantic state after it is reopened. We introduce Evidence-Guided Commitment Correction (EGCC), a training-free framework that treats factual revision as a controlled transition in this diffusion state space. Our key observation is that a dVLM can verify before it answers; at the still-masked answer state of a fixed object-existence probe, a continuous score from the dVLM’s native token distribution better ranks objects by unsupportedness than the probe’s hard decoded answer and remains informative even among commitments that receive the same decoded answer. Across all three dVLM backbones, unsupported object commitments also remain semantic attractors: simply reopening their spans reconstructs the same object in 80–86% of triggered cases. EGCC leverages masked-state evidence to detect these weak commitments and mediate a sparse transition to a locally reachable alternative. It preserves every non-target token by construction, constrains transition to be non-expansive over extracted object commitments, and requires each newly introduced object identified by the extractor to be better supported than the reopened source. Evidence estimation and revision use the same frozen dVLM, without an auxiliary verifier or full-response rollouts. With one numerical policy frozen across models and benchmarks, EGCC reduces measured object hallucination for LLaDA-V, MMaDA, and LaViDA on both COCO CHAIR and AMBER-Generative, with reductions of up to 3.8 CHAIR-S points and 11.7 AMBER Hal points. These results demonstrate masked-state evidence as an actionable interface for post-completion semantic correction in dVLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.