acceptodds
Under review as a conference paper at ICLR 2027

Right, but Can You Show It? Correct Answers from Multimodal LLMs Rarely Come with Faithful Evidence

Abstract

When a multimodal large language model (MLLM) answers a question about an image, it usually also states what it sees, such as which objects are present and where, as evidence for its answer. Users rely on this evidence to decide whether to trust the answer, yet benchmarks grade only the answer. They tacitly assume that *a correct answer comes with trustworthy evidence*. Testing this assumption means checking every stated fact against the image, and being able to prove a false fact false. That requires knowing everything the image contains: if a dataset labels only some of its objects, a claim about an unlabeled one might still be true and cannot be marked wrong, and most image datasets are labeled this way. Remote sensing is a rare exception. By protocol, its datasets label every object of declared categories, together with their spatial relations and their changes over time. Building on this, we introduce *evidence contracts*, which specify for each question the visual facts its answer depends on, and **RSFaith-Bench**, 13,511 human-validated questions over 16,288 images on which evidence is *checked against ground-truth annotation rather than judged by another model*. This check agrees with human raters on 89.4% of responses. Evaluating 22 MLLMs shows that the assumption fails in every one of them. Taken together, **more than half of the answers that accuracy counts as correct would fail a check of their own evidence.** Surprisingly, the dominant failure is not contradiction but missing support: most of these answers state nothing that the annotation refutes, yet their evidence does not establish the answer, a failure that no check of the stated claims alone can detect. More strikingly, **stronger models do not close the gap**. Across six larger or newer model versions, answer accuracy improves more than five times as much as faithful accuracy, and in half of them faithful accuracy *declines*. Human raters blind to our labels prefer the faithfully supported answer in up to 82% of matched pairs. Answer accuracy tells us when a model is right; it cannot tell us when to believe it. The [benchmark](https://huggingface.co/datasets/anonymous-paper-submit/RSFaith-Bench) and [code](https://anonymous.4open.science/r/RSFaith-Bench-submit-7CEE) are publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.