VEPO: Learning to Credit Visual Evidence Seeking for Agentic Video Understanding
Abstract
Long-video understanding requires acquiring sparse, temporally distributed visual evidence from long and redundant videos. Agentic approaches address this challenge through iterative temporal search, allowing models to gather visual observations before producing an answer. We find that correct answers produced by agentic systems are not always supported by the acquired evidence. During training with standard GRPO, answer accuracy improves while temporal evidence recall and active search decline, a training dynamic we characterize as evidence-seeking collapse. The objective rewards final-answer correctness, assigning identical rewards to evidence-supported predictions and unsupported guesses without checking whether sufficient evidence was acquired. Recent methods incorporate tool-use incentives and temporal localization rewards to encourage search activity and access to relevant regions, but these signals do not directly assess whether the accumulated observations provide sufficient evidence to answer the question. We introduce Video Evidence Policy Optimization (VEPO), which uses a frozen Evidence Reward Model to evaluate the sufficiency of accumulated observations. The resulting evidence signal allows VEPO to prefer evidence-supported answers over unsupported guesses and to credit the search decisions that improve evidence sufficiency. Across four long-video benchmarks, VEPO consistently outperforms supervised fine-tuning and standard reinforcement learning, improving both answer accuracy and grounded evidence acquisition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.