CoGA-GRPO: Correctness-Gated Reward Shaping for Video Reasoning
Abstract
Video reasoning requires balancing answer quality against generation cost. Uniform length incentives may reward unsuccessful responses, while fixed policy regularization fails to account for variations in observed success across rollout groups. We introduce CoGA-GRPO, a correctness-gated extension of group relative policy optimization. It uses a shared rollout-derived allocation score to adapt KL regularization and applies bounded, group-relative concision shaping only to eligible correct responses. Starting from the same reproduced structured-reasoning checkpoint based on Qwen2.5-Omni-7B, we train Standard GRPO and CoGA-GRPO under matched settings and evaluate both using the same four-frame protocol. Relative to Standard GRPO, the CoGA-GRPO accuracy point estimate increases from 66.60% to 67.53%, while the mean generation length decreases from 452.98 to 449.14 tokens. Component ablations and mechanism stress tests further characterize the observed accuracy–length trade-offs associated with the proposed design. CoGA-GRPO also achieves 59.82% accuracy on Daily-Omni. Code and evaluation artifacts are available in an anonymous repository at https://anonymous.4open.science/r/review-artifact-2027-81B5.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.