acceptodds
Under review as a conference paper at ICLR 2027

Beyond averaged policy gradients: RLVR learns dense parity from final rewards

Abstract

While reinforcement learning with verifiable rewards (RLVR) has achieved striking empirical success, a fundamental theoretical question remains open: when/how can final correctness alone provide an effective learning signal? In this work, we demonstrate that the limitation of ordinary Group Relative Policy Optimization (GRPO) in this setting is an optimization failure rather than an information deficit. Focusing on dense parity as a benchmark, we first prove that the population policy gradient underlying standard GRPO suffers from exponential decay, leaving an exponentially weak learning signal under the standard relative-reward mechanism. To resolve this, we introduce Verifier Posterior Transport (VPT), a sequential RLVR framework that executes a KL-regularized policy update after every example. Unlike GRPO, VPT guarantees a fixed, non-vanishing log-odds gain between successful and failed executions using only the exact same binary final reward. We construct a recursive transformer that realizes VPT exactly and prove that the example complexity required for target recovery scales linearly with input dimension at fixed confidence. Our results demonstrate that dense parity can be learned with linear sample complexity without intermediate supervision. Ultimately, this reveals that the bottleneck in final-reward RLVR is not whether the verifier holds sufficient signal, but how the update rule transforms that information into a learning signal.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.