Lost in the Zoom: Diagnosing and Rebalancing Evidence Use in Vision-Language Models
Abstract
Vision-language models (VLMs) with zoom-in tools improve visual question answering, yet acquiring useful visual evidence does not ensure correct reasoning over it. To understand this, we propose a new protocol to analyze tool-use trajectories and reasoning. The protocol separates trace curation, tool execution, evidence verification, and answer evaluation, then reuses trajectories for continuation comparisons. Applied to VLMs with zoom-in tools across six datasets, the protocol yields the *EvidenceZoom* benchmark, comprising 35.4K trajectories covering 3.9K questions. Attention diagnosis over trajectories reveals that reasoning consistently dominates attention allocation and that zoomed-in images are overlooked, with attention head patterns showing substantial variation in their focus within the zoomed-in region. Based on these findings, we propose *Head-Gated Attention Reallocation*, a training-free method that corrects the zoom–reasoning attention imbalance, scaling each head's correction by its relative within-zoom focus. Across five VLMs from the GLM, Qwen, and Gemma series, our method improves the six-dataset mean accuracy over the best training-free baseline and even surpasses the training-based method on Qwen3-VL-8B. To summarize, our work highlights the gap between acquiring visual evidence and using it effectively, and suggests attention reallocation as a promising direction for improving visual reasoning with zoom-in tools.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.