acceptodds
Under review as a conference paper at ICLR 2027

UNIVERSE: A Unified Benchmark for Video Deep Research

Abstract

Video deep research requires agents to coordinate video understanding with external information seeking, yet existing benchmarks cover only fragmented aspects of this capability and remain difficult to scale. Some provide a video and require research grounded in its content, while others require agents to discover relevant videos on the open web. Moreover, existing benchmarks largely emphasize using visual observations to initiate research, leaving the reverse process of using external knowledge to determine what should be observed underexplored. We introduce UNIVERSE, a unified benchmark that jointly evaluates supplied-video research and open-web video discovery. Beyond observation-to-research, UNIVERSE introduces knowledge-guided observation, where agents first resolve a target from indirect external clues and then locate it in video to obtain the requested visual evidence, together with image-based query clues that couple visual references with relational research. To enable scalable construction, we combine automated instance generation with multi-stage verification and shortcut filtering, yielding 800 high-quality questions across ten domains, substantially exceeding the scale of prior video deep research benchmarks. Evaluations of representative proprietary and open-weight multimodal models under a shared tool-enabled framework reveal substantial challenges in coordinating external research, video discovery, and visual observation, highlighting considerable room for progress in video deep research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.