Regulating What Policies Learn From: Geometry-aware Calibrated Policy Optimization
Abstract
Critic-free, group-based methods such as Group Relative Policy Optimization (GRPO) enable scalable LLM post-training from sampled rollouts, but typically do not account for differences in learning signal quality across rollout groups. Uncertainty provides a natural internal signal for regulating their contributions, yet existing rollout-level methods largely treat disagreement as a scalar proxy for reliability. We analyze rollout-level uncertainty from an optimization perspective and identify two key limitations through theoretical and statistical analysis: the anisotropic gap, where probability dispersion fails to capture the geometric magnitude of semantic disagreement, and the calibration gap, where uncertainty is decoupled from reward informativeness and therefore cannot distinguish noisy variation from useful reward-differentiated variation. To address these gaps, we propose Geometric-aware Calibrated Policy Optimization (GCPO), which combines geometry-aware uncertainty measures with reward-based calibration to regulate policy updates according to both disagreement structure and learning informativeness. Across QA and mathematical reasoning benchmarks, GCPO delivers stronger GRPO-style optimization than entropy-based uncertainty baselines, while statistical and ablation analyses support the underlying design assumptions. Our results suggest a broader principle for uncertainty-aware policy optimization: uncertainty should characterize not only how much rollouts disagree, but also whether that disagreement provides useful learning signals.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.