Getting the Evidence Right: Adapting Retrieval and Visual Reading for Multimodal Agentic Search
Abstract
Multimodal search agents must retrieve evidence that addresses intermediate questions and extract the facts needed to advance their reasoning. Existing approaches typically train the agent while keeping the retriever fixed. Retrieved candidates differ in how fully they support an intermediate query, and even relevant images can be misread. We introduce SPAR (Structuring and Projecting Evidence for Acquisition and Reading), a framework for adapting retrieval and visual reading. SPAR analyzes what each intermediate query requires and records the facts supported by each candidate. This shared analysis yields graded relevance supervision that prioritizes complete support while retaining useful partial evidence. For visual reading, SPAR reuses the image-grounded facts recorded in the same annotations to teach the agent to extract query-relevant information. We then adapt the agent's search policy on complete trajectories with the adapted retriever held fixed. On MultiModalQA, SPAR achieves 76.03% answer accuracy, exceeding the strongest baseline by 8.43 percentage points. Graded relevance supervision outperforms ungraded supervision both before and after policy adaptation. Under fixed evidence, the SPAR agent extracts visual facts more accurately than the policy-only agent. Without further training, SPAR also transfers to WebQA, MuSiQue, and HotpotQA, reaching 61.89% answer accuracy on HotpotQA. We will release our code upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.