acceptodds
Under review as a conference paper at ICLR 2027

EviGRAD: Evidence-Calibrated Rewards and Gradients for RLVR

Abstract

Verifiable-reward reinforcement learning (RLVR) for code generation relies on executable tests to provide training rewards. When these tests fail to cover important behavioral boundaries, incorrect programs may pass them and receive spuriously high rewards. We find an asymmetric reward-error pattern: high-reward false positives are relatively common and can be identified with useful accuracy, whereas low-reward candidates are rarely rescued by additional validated tests. This makes suspiciously high rewards the primary actionable source of reward misranking in group-relative policy optimization. We present , a risk-aware selective verification framework for conservative reward correction under a fixed verification budget. A lightweight reliability model ranks high-reward candidates and rollout groups according to their risk of being false positives. Only the highest-risk groups, together with a small exploration quota, receive additional verification and counterexample generation. Generated tests are admitted only after reference-based semantic validation and execution-based checks. Verified counterexamples then produce a one-sided reward correction: suspiciously high baseline rewards may be reduced, whereas low baseline rewards are not increased merely because a candidate passes an additional test. For rollout groups with degenerate low-signal rewards, EviGRAD uses dynamic resampling to recover learnable variation. The two mechanisms address complementary failure modes: erroneous positive training signals and insufficient within-group reward variation. We evaluate EviGRAD on CodeContests and APPS using matched models, training steps, rollout budgets, and verification costs. We measure agreement between training-time reward rankings and hidden-test rankings, the rate of false-positive positive-gradient signals, correction coverage, recovery of zero-advantage groups, final task success, and verification overhead. Across multiple random seeds and cost-controlled comparisons with binary GRPO, pass-rate GRPO, DAPO-style dynamic sampling, and always-augmented testing, EviGRAD reduces false-positive positive-gradient signals and improves the alignment between training-time rewards and hidden-test correctness, while dynamic resampling recovers learning signal from degenerate groups.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.