Selective Partition Repair for Group-Relative Policy Optimization
Abstract
Executable verifiers provide inexpensive rewards for visual-symbolic reasoning, but false negatives can eliminate useful training signals. In group-relative policy optimization (GRPO), an all-rejected group receives no ranking signal even when it contains correct responses. Correcting individual verdicts can still give a missed correct response a negative advantage, because each advantage depends on the rewards of the entire group. We formulate this problem as hidden-partition recovery and introduce Structured Group-Relative Correction (SGRC). SGRC learns a joint distribution over binary correctness masks from judge-annotated rollout groups. Its energy combines evidence from each response with interactions between the predicted correct and incorrect subsets. On subsequent groups, a confidence gate evaluates each complete predicted mask before applying it, leaving uncertain groups unchanged. The repaired rewards enter the original GRPO normalizer without changing the functional form of the policy objective. On the geometry sym subset of MathV360K, the complete SGRC system, using Qwen2.5-VL-3B-Instruct and 6,912 annotation attempts per run, achieves mean strict answer accuracy 2.96 percentage points above fixed-verifier GRPO across five paired training seeds. On held-out mixed groups, its joint predictor achieves exact-mask recovery 16.7 points above a parameter-matched factorized predictor. These results motivate evaluating group-relative reward repair by the complete partition it recovers, rather than only by individual verdict accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.