AUDIT-FIRST VAPO: RISK-CERTIFIED SELECTIVE UPDATES UNDER IMPERFECT VERIFICATION
Abstract
Verifier-guided reinforcement learning (RLVR) improves reasoning models by using automated evaluators to provide training signals, but imperfect verifiers can introduce harmful updates that degrade policy optimization. Existing approaches mainly optimize verifier quality or aggregate noisy feedback, without explicitly controlling the risk of accepted policy updates under limited verification budgets. We introduce Audit-First VAPO, a risk-certified selective update framework that decomposes verifier-guided optimization into update-direction certification and update-magnitude control. VAPO first audits candidate updates through sequential risk certification, selectively accepting, appealing, or abstaining from updates based on finite-sample guarantees, and then applies bounded policy updates only after certification. This design separates the decision of whether an update is safe from how strongly it should be applied, enabling robust learning under imperfect verification. We evaluate VAPO on two reasoning benchmarks with two language models against static RLVR, matched-random selection, confidence thresholding, noise-corrected RLVR, and verifier-augmentation baselines. At a target selected-risk level of , VAPO achieves 74.1% task accuracy with selected harmful-update risk of 0.0697, coverage of 0.4125, and relative verifier cost of . Under matched coverage and update magnitude, VAPO reduces selected risk by 0.0260 compared with matched random selection, with a paired 95% confidence interval of . Across asymmetric, confidence-dependent, and correlated verifier noise settings, the risk certificate remains satisfied in 57 of 60 independent runs. These results demonstrate that audit-first selective optimization provides a principled approach for reliable RLVR training when verification signals are costly and imperfect.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.