Re-Attending to the Missing Pieces: Test-Time Visual Evidence Refinement for High-Resolution Reasoning
Abstract
Multimodal large language models (MLLMs) can process increasingly high-resolution images, yet often miss small, question-relevant evidence in cluttered scenes. In our controlled study across three MLLMs, accuracy generally improves as the surrounding background is removed, while the queried object remains visible at every ratio. This exposes a visual evidence allocation bottleneck: more pixels do not ensure that finite encoding and attention budgets are directed to relevant evidence. Existing training-free methods may select or refine visual regions, but typically do not jointly diagnose missing evidence and verify the resulting crops. We introduce RECAP (Re-attending and Composing Attention-guided Proposals), a training-free closed-loop framework for adaptive test-time visual evidence refinement. RECAP identifies localizable entities and derives source-resolution proposals from their attention maps. Tentative answers direct refinement toward unresolved evidence, while comparative verification ranks proposals by target visibility. Selected regions form compact, order-aware canvases. Confidence-guided routing returns the answer once the model is sufficiently confident. Across V Bench, HRBench-4K, and HRBench-8K, RECAP improves all four evaluated backbones. It reaches 95.8% with Qwen2.5-VL-7B and 92.7% with InternVL3-8B on V*Bench, achieving the best results among the compared methods for both backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.