APEX: Accuracy Projection from Verifier-Labeled Experience for Offline RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) is often presented as online KL-regularized policy improvement, while practical offline pipelines cache responses and reuse verifier labels. After logging, the learner cannot explore beyond reference support, so those labels can only justify a reference-supported policy. Reward-weighted and rejection-sampling SFT convert verifier labels into imitation weights or filters, but they do not isolate the policy target induced by binary-verifier RLVR. We identify that target as the accuracy-optimal reference projection, which is the zero-temperature limit of KL-regularized RLVR and, equivalently, the minimum-KL reference-supported policy with unit verifier accuracy. This projection formulates offline RLVR as a Bernoulli density-ratio problem in which correct responses support the projected policy and incorrect responses identify reference mass to remove. (Accuracy Projection from Verifier-Labeled Experience) fits the KL reward-policy coordinate on logged reference rollouts with a binary correctness likelihood. Our analysis establishes population policy consistency, an exact accuracy-gap identity, a finite-sample bound scaling as for group size and correct-response fraction , and a bias-variance explanation for weighted and rejection-sampling SFT baselines. Ratio-stratified GSM8K studies and BigMath fine-tuning support this analysis, with the largest gains occurring in rollout groups that contain both correct-response support and incorrect-response labels. Beyond mathematical reasoning, also outperforms offline baselines and online GRPO on Search-R1 multi-turn search and question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.