acceptodds
Under review as a conference paper at ICLR 2027

Getting the Evidence Right: Adapting Retrieval and Visual Reading for Multimodal Agentic Search

Abstract

Multimodal search agents must retrieve evidence that addresses intermediate questions and extract the facts needed to advance their reasoning. Existing approaches typically train the agent while keeping the retriever fixed. Retrieved candidates differ in how fully they support an intermediate query, and even relevant images can be misread. We introduce SPAR (Structuring and Projecting Evidence for Acquisition and Reading), a framework for adapting retrieval and visual reading. SPAR analyzes what each intermediate query requires and records the facts supported by each candidate. This shared analysis yields graded relevance supervision that prioritizes complete support while retaining useful partial evidence. For visual reading, SPAR reuses the image-grounded facts recorded in the same annotations to teach the agent to extract query-relevant information. We then adapt the agent's search policy on complete trajectories with the adapted retriever held fixed. On MultiModalQA, SPAR achieves 76.03% answer accuracy, exceeding the strongest baseline by 8.43 percentage points. Graded relevance supervision outperforms ungraded supervision both before and after policy adaptation. Under fixed evidence, the SPAR agent extracts visual facts more accurately than the policy-only agent. Without further training, SPAR also transfers to WebQA, MuSiQue, and HotpotQA, reaching 61.89% answer accuracy on HotpotQA. We will release our code upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.