acceptodds
Under review as a conference paper at ICLR 2027

From What Is Seen to What Is Meant: How VLMs Infer Implication from Images

Abstract

Vision-language models can describe what an image depicts yet misinterpret what it means. This paper asks where that inference fails and examines four components: visual-evidence grounding, source-relation selection, target-domain mapping, and answer binding. We first conduct a four-stage audit of 975 incorrect II-Bench responses and assign 76.1% to source-relation selection or target-domain mapping, compared with 16.9% to explicit visual-evidence errors. Because response labels alone do not establish the computations behind these errors, we conduct controlled tests. Removing image content reduces accuracy, while three visual-attention summaries do not reliably distinguish correct from incorrect interpretations. When five true relations from the same image compete, the question-relevant relation ranks first in only 46.3% and 34.9% of full-context cases across the two models. Transferring relation-conditioned states changes both relation and answer preferences. With the correct relation fixed, implication-specific states steer preferences for held-out implications, and supplying an aligned implication improves final answer selection. These results show that the key bottleneck in image-meaning inference is not whether a model sees an image, but whether it can form and use the implication supported by the right relation among visible elements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.