Thinking Once is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA
Abstract
High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect images again through cropping, re-encoding, or multi-round search. We show that this view is incomplete: in many cases, fine-grained evidence has already survived visual encoding and become identifiable and influential within an intermediate-layer routing window, but is later diluted before answer generation. We propose Thinking-Once, a **training-free, single-visual-pass** evidence-routing method that reconstructs question-conditioned attention at this window, preserves core entity tokens and compact background context, and routes this evidence to later layers without extra visual encoding. Across five base models, Thinking-Once consistently improves or matches the corresponding base setting, increasing the average scores on Bench, HRBench-4K, and HRBench-8K by , , and points while reducing the average peak memory by about . On Qwen2.5-VL-7B, it improves the three benchmarks by , , and points, raising the cross-benchmark mean from to . With the ZwZ-8B base model, Thinking-Once reaches a mean score of . Against 11 open-source HR-VQA baselines, it obtains the best or tied-best score on all three benchmark averages and the best overall mean; for example, compared with DeepScan, it reduces Bench inference time by **%** while improving the cross-benchmark mean from to . These results show that HR-VQA can be improved by routing already encoded evidence rather than repeatedly acquiring new visual inputs. Code is available in the appendix.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.