MSearch: Benchmarking Multimodal Search Agents against Visual Misinformation
Abstract
Multimodal large language models have evolved into deep search agents that orchestrate retrieval tools and browse the open web for complex visual questions. Existing benchmarks, however, evaluate agents only under trustworthy evidence, measuring whether they can find the right answer but not whether they can withstand misinformation. We introduce MSearch, a benchmark that systematically examines multimodal search under AI-injected visual misinformation, constructing three levels—anchor entity replacement, context cue injection, and dual misinformation injection within and across images—via targeted image editing, yielding 330 human-audited samples. We evaluate 15 search agents spanning strong proprietary systems, lightweight models, and open-source multimodal search models. MSearch proves highly challenging: the average accuracy across all agents is merely 23.9%, with the strongest, Gemini-3.1-Pro-preview, reaching only 45.6%. Contrasting runs on pristine images with runs on misinformation-injected images reveals that misinformation substantially degrades performance: replacing only the image reduces accuracy by 19.2 points on average, and by as much as 28.8 points for the most susceptible agent. Trajectory analysis further reveals that injected misinformation steers agents toward erroneous search and analysis paths, with an average mislead rate of 42.5%. Agents also seldom spontaneously recognize conflicts between misleading visual cues and other available evidence: conflict awareness averages only 8.3%, and even the highest-performing agent on this measure, Gemini-3.1-Pro-preview, reaches only 25.0%. These findings underscore that, as image generation and editing models grow increasingly powerful, multimodal search agents must learn to verify visual cues against external evidence rather than comply with them. MSearch thus provides a rigorous testbed for diagnosing and defending against visual misinformation, paving the way toward trustworthy multimodal search.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.