acceptodds
Under review as a conference paper at ICLR 2027

WORTH: Decision-Aligned Credit Assignment for Evidence Acquisition in Long Video Agents

Abstract

Agentic video understanding has emerged as a promising paradigm for long-video reasoning, wherein an agent first perceives a sparsely sampled overview and then selectively retrieves evidence from task-relevant temporal segments. However, existing agents are typically optimized via outcome-based reinforcement learning with a single trajectory-level reward, which neither evaluates whether a given acquisition justifies its cost nor attributes the final outcome to individual observations. We term this deficiency the acquisition–attribution mismatch, and our systematic analysis shows that it drives agents to acquire evidence indiscriminately while most of the acquired observations contribute nothing to the final answer. To address this, we propose WORTH (Whether Observation Retrieval Truly Helps), a decision-aligned reinforcement learning framework that decouples acquisition from attribution via two dedicated credit signals. An acquisition credit estimates the ex-ante value of retrieval as the cost-adjusted gain of acquiring over answering immediately, computed through paired continuations from the same state. An observation credit estimates the ex-post contribution of each retrieved observation as its exact Shapley value, obtained by re-answering from every subset of acquired observations. Both credits are derived from the policy itself and replace the trajectory-level advantage exclusively on the tokens corresponding to their respective decisions. Experiments demonstrate that WORTH achieves state-of-the-art performance among most video agents across multiple long-video understanding and temporal grounding benchmarks, improving accuracy by 2.2% over outcome-based RL baselines while consuming 67% fewer tokens. Code, data, and models will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.