Lessons from Agentic Post-Training of MLLMs for Multimodal Web Search
Abstract
Multimodal web search requires models to formulate effective queries for knowledge-intensive visual questions, gather relevant evidence from the web, and synthesize grounded answers. We systematically study three aspects of post-training multimodal search agents: evaluation, supervised fine-tuning, and reinforcement learning. First, we quantify the impact of individual evaluation components on measured performance and advocate for a standardized evaluation protocol to enable reliable comparisons across methods. Next, we examine how the distribution of tool calls in the SFT training corpus shapes the student agent's search behavior. We further show empirically that short interaction horizons reduce computational costs while maintaining strong performance, whereas longer search horizons do not consistently improve accuracy. Finally, we investigate branching RL to refine intermediate search decisions using only terminal rewards. This approach generates alternative continuations from shared intermediate states and compares their terminal rewards, enabling more targeted credit assignment to individual search decisions. Our empirical analysis identifies early branching as an effective data-driven heuristic: concentrating branching on a small number of early search decisions yields more useful exploration than random branching and outperforms strategies guided by the model's own uncertainty, measured through token entropy or self-reported confidence. Extensive experiments show that the resulting RL policy improves average accuracy from 59.89% to 63.02% over its SFT initialization and achieves state-of-the-art performance on multiple benchmarks. Together, these findings provide practical guidance for evaluating and post-training multimodal search agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.