Auditing Verifier Cancellation in GRPO with Mixed Rewards
Abstract
Group-relative policy optimization (GRPO) subtracts each sampled group's mean reward. A binary verifier therefore contributes nothing to the advantages when every completion fails. A varying task reward can preserve nonzero advantages in the same group. We audit this known cancellation identity and examine its relationship to reward weighting and training recovery. Numerical replay reproduces task-only advantages in all 300 of 300 sampled all-violating groups after a constant verdict penalty is added, up to numerical rounding. We train a 0.5B policy on two fixed synthetic constrained tasks and compare reward designs and four remedy components at matched sampling budgets. The main comparisons omit KL regularization. Across the tasks' hard levels, a verdict below its crossover weight reaches final feasibility on 2 of 8 seeds. Task-only rewards reach it on 0 of 8 seeds. A verdict far above the crossover reaches it on 6 of 8 seeds. A graded residual, whose penalty also exceeds the crossover, reaches it on 7 of 8 seeds. A residual whose first-violation penalty equals the light verdict's reaches it on 7 of 8 seeds. That residual still orders every mixed group, and its penalty exceeds the light verdict's from the second violation unit. The stronger verdict acts through mixed groups that contain a feasible completion. Across the capacity task's hard-level reward and remedy arms, 23 of 24 feasible runs produce only parsed empty selections and yield no task value. We classify remedies by their inputs and show that filtering identical rewards retains every all-violating group with reward variation. Below the crossover on the speed task, the verifier filter and NGRPO with a feasible anchor reach final feasibility on 1 of 3 and 0 of 3 seeds. The identical-reward filter and Dr. GRPO reach it on 0 of 3 and 0 of 3 seeds. This is the predicted ordering, but it depends on that 1 verifier-filter seed. The remedy comparisons do not establish a stable efficacy ranking, and 0 of 50 capacity runs at weights below and above the crossover have outputs that are useful and feasible. The main experiments reuse training prompts for evaluation and test components within a simplified training loop. Our findings support reporting constraint satisfaction, task value, and output validity together. Acting on the cancelled groups does not settle whether training recovers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.