VeriSeek-VL: Seeking Visual Evidence and Verifying Hypotheses in Multimodal Deep Search
Abstract
Multimodal deep search agents solve long-horizon information-seeking tasks by iteratively retrieving and reasoning over textual and visual evidence. However, existing methods face two limitations: insufficient active acquisition of visual evidence and premature hypothesis lock-in. First, when key visual evidence is not explicitly provided, agents tend to rely on textual evidence and rarely actively acquire external visual evidence. Second, agents often treat an insufficiently supported hypothesis as fact in subsequent reasoning, so a wrong hypothesis leads the search away from the correct answer. To address these limitations, we propose VeriSeek-VL, a multimodal deep search framework that integrates data synthesis, training, and inference-time control. First, to strengthen active acquisition of visual evidence, we construct multimodal hypergraphs whose visual hyperedges link multiple entities through image-derived relations. Embedding these hyperedges into multi-hop reasoning chains yields tasks with implicit visual constraints, where key intermediate entities can be identified only through external visual evidence. Through supervised fine-tuning and GRPO on these tasks, the model learns to actively acquire external visual evidence. Second, to mitigate premature hypothesis lock-in, we identify a hypothesis lock-in axis through trajectory auditing and activation analysis. This axis is a linear direction in activation space that distinguishes lock-in steps from normal reasoning steps. During inference, we project each step's mean activation onto the axis and trigger verification of key hypotheses when the projection exceeds a threshold, reducing the risk of searching under a wrong hypothesis. Extensive experiments show that VeriSeek-VL achieves state-of-the-art performance among open-source methods on all five multimodal search benchmarks, surpassing the strongest baseline by 6.32% on average. Training improves the base model by 25.5%, and axis-guided verification adds a further 3.8%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.