acceptodds
Under review as a conference paper at ICLR 2027

AMR: Enhancing Visual Evidence Selection via Attention-Guided Multimodal Re-Reading

Abstract

Multimodal large language models (MLLMs) have advanced rapidly across vision-language tasks, yet their performance on visually grounded reasoning remains far from perfect. We investigate whether repeating the same image–question pair can address a limitation of causal, left-to-right processing, whereby visual evidence may be encountered before the model fully understands which information is relevant to the question. Repeating the complete multimodal input consistently improves inference, and an analysis of the repeated pass reveals a systematic redistribution of text-to-vision attention between its first and second reads. We refer to this read-to-read redistribution as *attention shift*. Rather than uniformly replaying the first allocation, the second read selectively increases attention to a subset of visual tokens, indicating that repetition can expose complementary visual evidence. Building on this observation, we propose *AMR* (Attention-Guided Multimodal Re-Reading), a training-free, plug-and-play intervention that leverages the model's own attention shift to identify and selectively reinforce visual evidence that becomes more relevant during the second read. Across six MLLMs and four multimodal reasoning benchmarks, AMR improves all 24 model–benchmark combinations over the zero-shot baseline, with an average gain of 2.61 percentage points. These results suggest that multimodal re-reading provides more than additional computation: it yields a state-dependent signal that can guide visual evidence selection at inference time, without parameter updates or task-specific training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.