UNIVERSE: A Unified Benchmark for Video Deep Research
Abstract
Video deep research requires agents to coordinate video understanding with external information seeking, yet existing benchmarks cover only fragmented aspects of this capability and remain difficult to scale. Some provide a video and require research grounded in its content, while others require agents to discover relevant videos on the open web. Moreover, existing benchmarks largely emphasize using visual observations to initiate research, leaving the reverse process of using external knowledge to determine what should be observed underexplored. We introduce UNIVERSE, a unified benchmark that jointly evaluates supplied-video research and open-web video discovery. Beyond observation-to-research, UNIVERSE introduces knowledge-guided observation, where agents first resolve a target from indirect external clues and then locate it in video to obtain the requested visual evidence, together with image-based query clues that couple visual references with relational research. To enable scalable construction, we combine automated instance generation with multi-stage verification and shortcut filtering, yielding 800 high-quality questions across ten domains, substantially exceeding the scale of prior video deep research benchmarks. Evaluations of representative proprietary and open-weight multimodal models under a shared tool-enabled framework reveal substantial challenges in coordinating external research, video discovery, and visual observation, highlighting considerable room for progress in video deep research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.