Calibrated Bernoulli Policy Optimization for Single-Sample RLVR
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as the core paradigm for enhancing reasoning capabilities in Large Language Models (LLMs). Currently, multi-sample algorithms like GRPO dominate this field despite their prohibitive computational costs. While single sample sequence level methods offer a theoretically efficient alternative, they suffer from notorious training instability. This paper reveals the mechanism underlying the empirical robustness of GRPO: at the infinite sample limit, it implicitly approximates Bernoulli variance normalization. However, due to the inherent calibration deficit of critic networks, direct introduction of this variance normalization triggers training instability. To address this, we propose Calibrated Bernoulli Policy Optimization (CBPO). By integrating Expected Calibration Error (ECE) regularization and lightweight L-BFGS temperature scaling, CBPO aligns the critic probability estimations. This calibration enables stable analytical variance normalization, yielding Fisher information aware advantage scaling without the need for multi-sampling. Extensive experiments demonstrate that CBPO matches the training stability and SOTA performance of multi-sample baselines, requiring only a single sample computational budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.