acceptodds
Under review as a conference paper at ICLR 2027

When Rewards Disagree: Reward Weighting for Token-Efficient Video Reasoning

Abstract

Human perception interprets video dynamics by constructing structured mental representations of entities, actions, and temporal boundaries before engaging in high-level reasoning. In contrast, current Video-LLMs typically rely on unstructured Chain-of-Thought (CoT), where crucial visual cues are obscured by verbose narration. While recent efforts ground reasoning in intermediate visual evidence, they rely on task-specific schemas that fail to generalize across diverse video understanding tasks. To address this, we introduce Structured Event Evidence, a unified, temporally ordered schema designed for general video reasoning. However, post-training a model to extract structured evidence while reasoning introduces a severe multi-objective optimization bottleneck, as candidate rollouts must simultaneously satisfy schema compliance, task accuracy, and reasoning budget. Standard reinforcement learning (RL) algorithms attempt to balance these objectives via static scalarization, enforcing a rigid accuracy–length trade-off across training and obscuring fine-grained reward signals. We resolve this by proposing Group-Relative Advantage Balancing (GRAB), which dynamically computes Pareto-optimal objective weights within each rollout group via a minimum-norm criterion over centered rewards. Operating strictly in reward space, GRAB incurs virtually zero computational overhead (<0.002% relative runtime to GRPO) and eliminates brittle weight tuning. Supported by our new PARSE-60K dataset (60K annotations across 35K videos) and a three-stage progressive curriculum, we train PARSE-4B. Across temporal grounding and video question-answering benchmarks, GRAB generates 13% to 39% fewer tokens while maintaining competitive accuracy with exhaustively hand-tuned static weights—without requiring hyperparameter weight sweeps.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.