acceptodds
Under review as a conference paper at ICLR 2027

Necessity-Bench: A Self-Report-Free Test of Whether Vision-Language Model Answers Are Actually Grounded in Evidence

Abstract

Vision-language models (VLMs) often answer questions about images correctly, but it is unclear whether those answers actually depend on the part of the image that matters. Existing ways of checking this rely on the model's own account of its reasoning: they either reward overlap between a region the model cites and the correct object, or test whether the model's written rationale mentions an object that was removed. Both trust self-report, which can be wrong. We introduce Necessity-Bench, a protocol that uses no self-report at all. For each question, an independent oracle (a scene program or scene graph, never the model) identifies the region the answer depends on and a separate, irrelevant control region. We edit each region in turn and ask the model again, checking three properties of its answer: sufficiency (the relevant region alone reproduces the answer), necessity (removing it changes the answer to the true counterfactual answer, recomputed by re-running the ground-truth program), and specificity (removing the irrelevant region leaves the answer unchanged). The control rules out models that change their answer under any edit, and the recomputed counterfactual is stricter than testing for "any change." The protocol requires no training. On CLEVR (3,000 items; 853 with a well-defined counterfactual), Qwen3-VL-8B scores 0.680 on the strict joint metric versus 0.619 for Qwen3-VL-2B (p=0.0097), a scale effect driven by correctly tracking the counterfactual rather than by answer stability. On real photographs (GQA, 962 items), specificity stays above 0.92 across 2B/8B/32B, but no scale effect appears (p≥0.264). Finally, we ask whether self-report is safe to use as a training signal. Reinforcement learning on a model's own predicted evidence box raises every self-reported metric (2B accuracy 0.617→0.889; 8B box IoU 0.295→0.597). Yet our protocol, run on items excluded from training, shows that the 2B model's true causal grounding falls (necessity 0.656→0.537, p=2.9×10⁻⁶) while the 8B model is unchanged. Self-reported progress and real grounding can move in opposite directions: grounding should be verified by intervention, not by asking the model. Necessity-Bench thus offers a training-free, self-report-free measure of grounding that exposes failures self-report conceals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.