acceptodds
Under review as a conference paper at ICLR 2027

Vision-AutonoSearch: Learning to Investigate Images

Abstract

Multimodal search agents answer knowledge-intensive visual questions through multi-turn search, but the question typically already specifies the search direction. Many images, such as memes, advertisements, artworks, and news photographs, instead derive meaning from the context they reference or imply, so explaining them requires deciding what to investigate. We formulate this task as Visual Autonomous Investigation (VAI): given an image and a general explanation request, the agent decides what to investigate, gathers evidence, and produces a grounded explanation. We present Vision-AutonoSearch (VAS), a framework that trains a single policy for VAI and search-augmented VQA through supervised fine-tuning and reinforcement learning. We build VAI-SFT-5K with investigation demonstrations and VAI-RL-2K with private rubrics and reference evidence. Rubrics score essential content but cannot distinguish explanations that meet the same criteria yet differ in accuracy, evidential support, or additional findings. We therefore propose Evidence-Constrained Comparative Reward (ECCR), which complements rubric scores with pairwise comparisons among sampled investigations. A preference is granted only when visual or retrieved evidence substantiates its underlying claims, preventing persuasive but unsupported content from being rewarded. Across four VAI and four VQA benchmarks, including our VAIBench, Vision-AutonoSearch-9B achieves the highest VAI average (66.2) and overall average (61.4) among the compared models, and exceeds three specialized search agents on both task averages. Ablations show that ECCR raises VAIBench from 69.9 to 73.6 and reduces the unsupported claim rate from 10.6% to 7.3%. Adding VAI data to VQA training under a matched budget improves all three VQA benchmarks, suggesting that learning to investigate strengthens capabilities shared with question answering.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.