acceptodds
Under review as a conference paper at ICLR 2027

Reinforcing Video Reasoning via Fusion-Aware Chain-of-Thought

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has recently been extended to video reasoning, where perceptual grounding must be integrated with reasoning across space and time. However, Group Relative Policy Optimization propagates a single outcome-based advantage across all tokens, assigning similar credit to decision-promoting and decision-undirected information within Chain-of-Thought. Our analysis reveals a fundamental contrast: whereas image–text RLVR concentrates probability mass and progressively reduces trajectory entropy, video sustains correct trajectories across high probability–entropy bands. This observation suggests that video reasoning benefits from preserving uncertainty during temporal evidence accumulation and reducing it during reasoning. We therefore propose erception–easoning usion-ware olicy ptimization , which derives a perception–reasoning fusion gap from probability–entropy statistics over perceptual-prefix and reasoning-suffix segments. PR-FAPO maps this signal to token-wise discount coefficients, enabling fusion-aware credit assignment that promotes informative early perception and selectively fuses articulated visual information into subsequent reasoning. Experiments on reasoning-centric and temporal-grounding benchmarks validate the effectiveness of PR-FAPO, demonstrating performance improvements together with improved training stability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.