acceptodds
Under review as a conference paper at ICLR 2027

BrowseComp-V³: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents

Abstract

Multimodal large language models (MLLMs) are evolving into browsing agents capable of multimodal deep search, supported by advances in autonomous planning and tool use. However, existing benchmarks remain limited in task complexity, information searchability, and evaluation dimensions, hindering comprehensive assessment in open-world settings. To address these limitations, we introduce BrowseComp-V³, a benchmark comprising 300 challenging, manually constructed questions across diverse domains. The tasks emphasize multi-hop, cross-modal reasoning at multiple levels of complexity and are designed to require evidence gathering through web browsing. We verify the public searchability of key evidence and provide expert-validated subgoals for process evaluation, enabling fine-grained analysis of search behavior and model capabilities. We also provide OmniSeeker, a general multimodal browsing agent framework, and evaluate a broad range of MLLMs and browsing systems. Even OpenAI Deep Research, the best-performing proprietary system in our evaluation, achieves a success rate of only 56% on BrowseComp-V³. Further analysis reveals bottlenecks in multimodal information integration and fine-grained perception, suggesting that effectively using visual evidence in multi-step search and reasoning remains challenging.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.