Video-SR1: Reinforcement Learning with Self-Rewarding Video-Language Reasoning
Abstract
Vision-language models increasingly support video, but reinforcement learning (RL) for video reasoning still relies largely on sparse final-answer rewards. Such supervision is poorly suited for long video-reasoning traces, where correct observations and localized perceptual or temporal errors may coexist within the same response. We introduce Video-SR1, a self-rewarding framework for dense RL post-training of video-language models. Its core algorithm, Group and Sentence Relative Policy Optimization (GSRPO), uses the policy to apply self-rewards across each sentence in its own reasoning trace, aligns semantically corresponding sentences across responses, and performs group-relative credit assignment at the sentence level. Further, GSRPO complements the self-rewards with a counterfactual temporal-sensitivity reward that encourages reasoning to respond to changes in video frame order. Across six diverse video-reasoning benchmarks, Video-SR1 significantly improves accuracy, sample efficiency, and cross-domain transfer relative to strong baselines in settings with and without the sparse ground-truth answer reward included during RL. Error analyses and ablations reveal two complementary factors underlying Video-SR1's gains: the temporal-sensitivity reward promotes reasoning that responds to the video's spatiotemporal structure, while local self-rewards ensure the temporally-sensitive reasoning remains accurate and grounded in the visual evidence. Together, our results suggest that a central limitation of RL for video reasoning is not only reward quality, but reward resolution: assigning credit at the level of intermediate claims can make each rollout more informative and provide a stronger training signal for grounded multimodal reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.