acceptodds
Under review as a conference paper at ICLR 2027

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Abstract

Open-world video understanding often requires locating sparse visual evidence in a video and acquiring knowledge that is contained neither in the video itself nor in the model's parameters. Thinking-with-Videos actively localizes such evidence, and Deep Research acquires such knowledge over multiple search steps, but the two directions are developed in isolation. To bridge this gap, we introduce VideoRover, a unified Video Deep Research framework that iteratively coordinates video cropping, multimodal search, and webpage browsing. Given a video-question pair, VideoRover conditions each action on the observations accumulated so far, so that localized clips anchor external search, while retrieved evidence triggers further video inspection and verification. Existing data rarely cover such trajectories, so we build an automated pipeline that constructs questions jointly dependent on video and web evidence and synthesizes verified interaction trajectories, yielding 26K SFT examples and 3K hard RL instances. We also introduce VideoRover-Bench, a benchmark stratified by video duration and research difficulty. Experiments on VideoDR and VideoRover-Bench show VideoRover-8B-RL performs on par with proprietary models in direct-answer settings while outperforming some larger models given the same tools. Additional experiments further demonstrate broad generalization and adaptive coordination between video grounding and external retrieval.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.