acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Look: Attention-Proposed Temporal Rewards for Grounded Video Reasoning

Abstract

Video reasoning is a core capability of multimodal large language models (MLLMs), crucial for downstream tasks such as video question answering and captioning. Recent work has advanced this capability through reinforcement learning (RL), following its success in the language domain. However, video reasoning poses distinct challenges. A model must first locate the segments relevant to the question along the temporal axis, then reason from the visual evidence they provide. Existing rewards supervise only the final answer, leaving the first of these unconstrained. Recent approaches annotate temporal spans or invoke external perception tools at inference, incurring annotation or computational cost, and even an explicit interval reward leaves the fraction of correct answers unsupported by the model's own attention unchanged. We propose VideoAPR, a reinforcement learning framework that grounds video reasoning in the model's own attention. We introduce AP-GRPO, which derives a temporal reward from where the policy already attends, requiring no temporal annotation during training. Within each rollout, position-debiased attention proposes a candidate temporal segment, and the response the policy has already produced is rescored with the visual context restricted to that segment and to an equal-length control window elsewhere in the clip, with the margin between the two forming the reward. Verification is performed only during training, so inference remains a single forward pass at no additional cost. Extensive experiments demonstrate that VideoAPR achieves state-of-the-art performance across diverse video reasoning benchmarks, outperforming recent methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.