acceptodds
Under review as a conference paper at ICLR 2027

Success-aware Policy Learning in RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) remains sample-inefficient on hard prompts, where most rollouts fail and the learner receives only a binary correctness signal. In the standard sample–verify–update loop, a rare successful rollout is used primarily as a target for parameter updates, leaving its reasoning content unused for acquiring subsequent training experience. We introduce Success-Aware Policy Learning (SAPL), which gives each verified success a second role of temporary in-context guidance for further sampling before the policy is updated. SAPL first discovers verifier-correct on-policy rollouts and selects the success with the highest likelihood under the current policy. The same model, with unchanged parameters, then conditions on this success to generate additional candidates. The verifier filters these candidates, and a weighted objective assimilates the verified responses into the prompt-only policy. SAPL requires no reference solution, external teacher, or verifier feedback beyond correctness, and uses no auxiliary context at inference. We show that SAPL achieves improved accuracy across Qwen3-4B, 8B, and 14B on eleven competition mathematics benchmarks by – percentage points over the strongest evaluated RLVR baseline and by – points over the strongest evaluated ground-truth-conditioned baseline. These results show that rare successes are valuable not only as update targets, but also as guides for acquiring new verified experience.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.