Counterfactual Auditing of Object Presence and Quantity Judgments in Vision-Language Models
Abstract
Vision-language models (VLMs) increasingly support systems that interpret visual content, making reliable evaluation essential. A basic capability is judging object presence, yet a correct answer on a static benchmark does not establish whether it relies on the relevant visual evidence. We introduce Counterfactual Object Visual Evidence (COVE), a paired-image benchmark that tests whether models update their answers when supporting object evidence is removed. Existence edits aim to remove the queried category, whereas quantity edits leave one instance, preserving category evidence while invalidating the at-least-two condition. The core datasets contain 160 existence pairs and 1,000 quantity pairs, complemented by 906 irrelevant-region existence controls. COVE measures evidence persistence: affirmative answers on both the original and target-edited image, stratified by intervention quality. Across six open VLMs, all-pair persistence ranges from 2.9–35.3% on 102 visually screened existence pairs and from 50.2–95.1% on a separate cohort of 329 metadata-screened two-to-one quantity pairs. More than 97% of eligible original-affirmative responses remain affirmative under irrelevant existence edits. Substantial count persistence recurs across prompts and editors; separate visual-audit aggregates provide a descriptive cross-check. These profiles show that low existence persistence can coexist with substantial difficulty updating quantity judgments. By combining instance-level removal records, quality strata, and paired response metrics, COVE provides a reusable resource for evaluation beyond static correctness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.