acceptodds
Under review as a conference paper at ICLR 2027

Prior Locking in RLVR: Prior-Corrected Optimization for Broader Reasoning Coverage

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving reasoning in large language models (LLMs). However, increasing aggregate success probability does not necessarily imply broad coverage over distinct correct reasoning modes. We show that, under the common binary-reward setting, even idealized KL-regularized policy improvement can exhibit prior locking, whereby the allocation among correct reasoning modes remains tied to initialization and limits coverage. To overcome this limitation, we propose Prior-Corrected RLVR (PC-RLVR), which dynamically rebalances inherited probability allocation; its oracle dynamics converge to the coverage-optimal uniform allocation over correct modes. To bridge the gap between the mode-level oracle and practical LLM training, we develop a historical-reference-based response-level prior correction with bounded weights, yielding a more stable practical realization under limited rollout budgets. We further derive a finite-time stationarity bound separating structural approximation residuals from controllable optimization and historical-reference estimation errors, with the controllable terms decaying as . Experiments on mathematical reasoning benchmarks show improved Avg@64 and Pass@64, together with stronger log-likelihood gains for initially lower-prior correct responses, providing complementary evidence consistent with broader coverage of successful responses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.