acceptodds
Under review as a conference paper at ICLR 2027

When Learning Misses the Target: Bridging the Gap Between RLVR and its Optimum

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning ability of large language models using outcome-based signals.However, it remains unclear whether policies learned by practical, sample-based RLVR recover the optimum of the underlying KL-regularized objective. To explore this, we derive the closed form solution of the desired optima, i.e., a distribution of reference policy reweighted by a Gibbs distribution, whereas the Gibbs term pushes policy to explore on tokens with higher expected reward. We further find that practical RLVR policies deviate substantially from this structure and concentrate probability on a narrower set of successful solutions.We therefore introduce Reference-Anchored Policy Correction (RAPC), which freezes the reference language model and learns a bounded process-level correction from terminal verifier feedback. The resulting acting policy combines the reference distribution with the learned reward-dependent correction. On reasoning benchmarks, RAPC more closely follows the predicted Gibbs structure and retains greater coverage of successful paths than updating the full language model under comparable evaluation protocols.These results suggest that practical RLVR can better exploit the information encoded in the KL-regularized optimum by explicitly learning its reward-induced correction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.