acceptodds
Under review as a conference paper at ICLR 2027

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Abstract

How reliably can we expect a multimodal large language model (MLLM) to use or disregard an image that conflicts with prior knowledge? We introduce the WhatIfVis benchmark, adapting and extending counterfactual datasets across five coarse-grained conflict types, such as color and size. Each example holds the image and question fixed while changing only the instruction to use or ignore the image. Across six MLLMs, mean paired accuracy is only : success requires ranking the intended answer above the alternative under both instructions. Yet this failure is ambiguous: visual evidence may be missing from the model's representations (H1), or remain available but be mishandled (H2). Reconstruction models trained without counterfactual examples recover coarse attributes from the final-layer image tokens of three frozen MLLMs, arguing against H1 for these attributes. LoRA adaptation of the language backbone on one conflict type, with the vision encoder frozen, raises mean paired accuracy to , with gains extending to four held-out types. Activation patching identifies layer windows where interventions change whether answers follow the image or prior knowledge. A learned rank-1 intervention transfers from each adapted model to its original, adapter-free counterpart, reaching mean paired accuracy without verbal intent instructions. Matched text statements make the same task-relevant facts easier to follow when instructed, providing a comparison that also reflects differences in evidence extraction. MLLMs can retain visual evidence yet fail to control its use. WhatIfVis exposes this gap and provides a testbed for developing more robust multimodal models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.