acceptodds
Under review as a conference paper at ICLR 2027

Calibrated Bernoulli Policy Optimization for Single-Sample RLVR

Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as the core paradigm for enhancing reasoning capabilities in Large Language Models (LLMs). Currently, multi-sample algorithms like GRPO dominate this field despite their prohibitive computational costs. While single sample sequence level methods offer a theoretically efficient alternative, they suffer from notorious training instability. This paper reveals the mechanism underlying the empirical robustness of GRPO: at the infinite sample limit, it implicitly approximates Bernoulli variance normalization. However, due to the inherent calibration deficit of critic networks, direct introduction of this variance normalization triggers training instability. To address this, we propose Calibrated Bernoulli Policy Optimization (CBPO). By integrating Expected Calibration Error (ECE) regularization and lightweight L-BFGS temperature scaling, CBPO aligns the critic probability estimations. This calibration enables stable analytical variance normalization, yielding Fisher information aware advantage scaling without the need for multi-sampling. Extensive experiments demonstrate that CBPO matches the training stability and SOTA performance of multi-sample baselines, requiring only a single sample computational budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.