acceptodds
Under review as a conference paper at ICLR 2027

Believing Is Seeing: How VLMs Assemble Trust in an Image, and Where It Can Be Read

Abstract

When an image contradicts what a vision-language model expects, the model has to decide whether to believe it. We study how CLIP-based LLaVA models make that decision, using image–prior conflicts in which we can edit, shrink, or segment the contradicting object and intervene on the model's attention and residual stream. We find that the decision is read from a single representation at the answer position, and that three very different interventions move it in the same way: degrading the contradicting object, shrinking it, and redirecting a small set of attention heads toward it. A 256-dimensional subspace of the residual stream is enough to carry most of the effect of an image edit, and it generalizes to conflicts it was not fit on, although the model does not depend on it exclusively: removing it changes only a fraction of decisions. Layer-wise patching locates the image's contribution in the image tokens through the first third of the network and at the answer position from the middle onward, where the attention heads that mediate the decision act. The same map appears on a second encoder family (Qwen2.5-VL), and the readout and the edit sensitivity are already present in a LLaVA-1.5 that was never instruction-tuned on images. We then test the readout on real photographs from NaturalBench and find that the decision is again formed by the middle of the network, that an untrained mid-depth readout predicts it reliably, and that most of what is being overridden there is the model's default yes/no answer rather than knowledge about the object. The readout is accurate enough to let the model stop early on the items it is confident about, but the interventions that move trust on curated conflicts do not correct it on photographs. Together, these results describe trust in the image as one representation with a known content, location, and depth, and show that the curated-conflict results do not predict behavior on natural images beyond the readout.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.