acceptodds
Under review as a conference paper at ICLR 2027

Teaching Video Agents How to Look via Counterfactual Information Gain Policy Optimization

Abstract

Video agents are increasingly equipped with tools that actively retrieve visual evidence during reasoning. However, video retrieval can be costly, as each retrieved video evidence can introduce thousands of context tokens, and noisy, irrelevant frames can hinder model performance. In this work, we focus on developing efficient video agents that make intelligent tool decisions about when to look (whether calling a tool is necessary) and where to look (what video interval to retrieve to maximize relevant information). Existing reinforcement learning frameworks for video agents typically use outcome rewards or dense turn-level rewards, which encourage spurious tool calls and struggle to learn stopping conditions and temporal localization. To address this challenge, we introduce Counterfactual Information Gain Policy Optimization (CIGO), an RL framework that incentivizes tool calls based on the model’s intrinsic information gain compared to directly answering the question. To accurately attribute information gain to the tool calls, we sample two types of counterfactual actions branching from the same reasoning prefix: (1) thinking and answering directly, and (2) issuing a single tool call before answering. By using this counterfactual construction, the model learns the stopping condition, while variance within the tool-call subgroup grounds localization. Extensive experiments across multiple benchmarks demonstrate improved overall performance on complex video reasoning tasks compared to existing agents. Beyond aggregate performance, CIGO learns more adaptive tool use and retrieves higher-quality visual evidence than existing video agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.