RevMMSearch: Evidence-Grounded Identity Revision for Multimodal Search Agents
Abstract
Multimodal search agents answer questions grounded in images by combining visual understanding with external retrieval and reasoning. However, visual understanding can be unreliable when resolving target identity, as mistaking the target for a visually similar entity can steer subsequent retrieval and reasoning along an incorrect trajectory. In this work, we argue that target identity should be treated as a revisable routing state, with search both conditioned on the current identity hypothesis and used to refine it. Our pilot studies further show that such errors can persist through stale evidence, constrain subsequent search, and complicate outcome-based credit assignment. Motivated by these findings, we introduce RevMMSearch, a framework that maintains competing target-identity hypotheses and actively acquires external evidence to verify and revise them throughout search. Within this framework, an entity-evidence graph links identity hypotheses to supporting or conflicting evidence and identity-dependent answer facts, dynamically updating fact applicability as the identity is revised. A unified search loop interleaves answer-fact acquisition with discriminative searches that distinguish competing hypotheses, while verifying whether retrieved evidence is attributable to the visual target. We further introduce revision-aware policy optimization, which combines final-answer rewards with evidence-based feedback on intermediate decisions and suppresses outcome reward signals that conflict with this feedback. Experiments on five benchmarks show that RevMMSearch improves average answer accuracy by 3.5% over the strongest compared multimodal search agent. Ablations and further analyses attribute these gains to reduced identity lock-in, filtering of stale evidence, and conflict-aware credit assignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.