Reinforcing Video Reasoning via Fusion-Aware Chain-of-Thought
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has recently been extended to video reasoning, where perceptual grounding must be integrated with reasoning across space and time. However, Group Relative Policy Optimization propagates a single outcome-based advantage across all tokens, assigning similar credit to decision-promoting and decision-undirected information within Chain-of-Thought. Our analysis reveals a fundamental contrast: whereas image–text RLVR concentrates probability mass and progressively reduces trajectory entropy, video sustains correct trajectories across high probability–entropy bands. This observation suggests that video reasoning benefits from preserving uncertainty during temporal evidence accumulation and reducing it during reasoning. We therefore propose erception–easoning usion-ware olicy ptimization , which derives a perception–reasoning fusion gap from probability–entropy statistics over perceptual-prefix and reasoning-suffix segments. PR-FAPO maps this signal to token-wise discount coefficients, enabling fusion-aware credit assignment that promotes informative early perception and selectively fuses articulated visual information into subsequent reasoning. Experiments on reasoning-centric and temporal-grounding benchmarks validate the effectiveness of PR-FAPO, demonstrating performance improvements together with improved training stability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.