acceptodds
Under review as a conference paper at ICLR 2027

EViC-VL: Evidence-Conditioned Verification for Confident Object Hallucinations in Vision-Language Models

Abstract

Vision–language models (VLMs) can produce object claims that are both linguistically confident and unsupported by the image. Such claims can evade verification policies that intervene only when a model exposes high self-uncertainty. We study this failure mode as confident visual hallucination (CVH). We operationalize CVH at the object-claim level and construct ConfHall-VL, a COCO-derived diagnostic that labels each candidate purely from the caption's own confidence and uncertainty signals together with the image annotations, without any detector evidence. We then propose EViC-VL, an evidence-conditioned procedure that queries two frozen open-vocabulary object detectors, OWLv2 and Grounding DINO, with canonical object names, averages their evidence scores, and applies a validation-frozen policy to the original caption. As a secondary robustness check, EViC-VL matches the frozen base generator on standard hallucination benchmarks overall, with its largest gain on object-presence polling accuracy (POPE F1). Our central question is whether combining these independent detectors improves object-level CVH discrimination over a single detector. On a sealed holdout, the equal-weight ensemble reliably raises the area under the ROC curve (AUROC) over the single-detector baseline by about four points. Overall, EViC-VL treats detector output as fallible evidence rather than ground truth, and this design measurably improves confident-hallucination detection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.