acceptodds
Under review as a conference paper at ICLR 2027

VideoSeeker: Empowering Cross-Video Understanding via Multi-Turn Evidence-Seeking Agentic Reasoning

Abstract

Cross-video understanding (CVU) requires models to locate key evidence across videos, establish temporal correspondences, and integrate intermediate findings through explicit comparison. Evidence needs become clearer as reasoning unfolds, yet end-to-end methods that answer directly from fixed frame samples lack mechanisms to supplement and verify evidence based on intermediate judgments. Under limited context budgets, sparse sampling may miss brief yet critical events, while dense encoding of entire videos scales poorly to multiple long videos. To this end, we formulate CVU as multi-turn agentic reasoning driven by active evidence seeking, integrating targeted segment replay, temporally interleaved comparison, and persistent evidence memory into a closed loop of observation, comparison, and reasoning. To learn reliable evidence-seeking policies, we first construct 4,000 multi-turn trajectories through verifiable trajectory distillation for supervised fine-tuning. During reinforcement learning, we assign fine-grained credit to evidence notes using information-gain rewards estimated from changes in the log-likelihood of the GT under the current policy. We further introduce a visual grounding gate that restricts positive credit to notes citing previously observed frames, mitigating reward hacking. Extensive experiments demonstrate the effectiveness of our multi-turn agentic reasoning framework for CVU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.