acceptodds
Under review as a conference paper at ICLR 2027

Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom

Abstract

High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce *VPS*, a visual parallel-search framework in which a main agent first invokes *grid_search* to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes *zoom_in* on a precise or merged region. The same interface supports both training-free inference and post-training of the main and sub-agents. Across five benchmark splits and three model scales, *VPS* improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to points and especially strong improvements for smaller main models. ZoomBench retains an approximately -point gain at every tested scale. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO objective for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a -point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from to while achieving higher pass@1 on most benchmarks, and joint training reveals a sharp asymmetry between local evidence reading and global search control. Together, these results establish *VPS* as both an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.