When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes
Abstract
Visual question-answering systems should answer when evidence is sufficient and abstain when it is not. Evaluating this behavior requires justified labels, because removing image content can leave the answer recoverable. We introduce a benchmark spanning PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each question groups original and edited images, presented independently; success requires correct answers on all supported versions and abstention on all versions labelled unanswerable. Complementing human or model validation, chart missing-information labels include an executable check: two complete charts with permitted values give different answers but identical pixels after masking. This demonstrates missing information under the specified chart rules. CLEVR and GQA labels instead follow programs operating on scene descriptions and designated edits; photographic masks are analyzed for residual location cues without the same ambiguity guarantee. Across 72,000 responses from six model configurations, the highest observed group success rates are 57.0%, 43.5%, and 33.7%, respectively. The benchmark separates failures in deciding whether to answer from failures to supply the correct answer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.