acceptodds
Under review as a conference paper at ICLR 2027

Knowing Where Without Reading What: Saliency-Induced Access Failure in MLLMs

Abstract

Visual question answering (VQA) requires models not only to locate task-relevant evidence but also to use it despite more salient competitors. Although multimodal large language models (MLLMs) excel at general VQA, their ability to exploit correctly localized evidence against salient distractors remains largely unexplored. To address this gap, we present **DECOY**, a benchmark designed to systematically evaluate MLLMs under conflicts between visual saliency and task relevance. DECOY contains 2,163 images from four eye-tracking datasets and 9,596 human-written question-answer pairs across five task types, each graded by a fixation-based saliency gap. Separate evaluations of *Saliency Grounding*, *Target Grounding*, and *Answer Accuracy* isolate suppression failures after correct grounding. Experiments on 24 MLLMs reveal that models know what is salient and where the evidence is, yet let the salient distractor dominate the answer: the lowest suppression failure rate is 37.7%, versus 4.9% for humans, and answering errors increase monotonically with the saliency gap. Beyond benchmarking, we uncover a gap between evidence localization and utilization. Restricting the input to target evidence is substantially more effective than indicating its location, and answer-relevant information remains decodable even as the causal influence of visual-region states fades with depth. These findings establish DECOY as a testbed for improving evidence utilization. The benchmark and code will be released upon publication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.