acceptodds
Under review as a conference paper at ICLR 2027

Regulating What Policies Learn From: Geometry-aware Calibrated Policy Optimization

Abstract

Critic-free, group-based methods such as Group Relative Policy Optimization (GRPO) enable scalable LLM post-training from sampled rollouts, but typically do not account for differences in learning signal quality across rollout groups. Uncertainty provides a natural internal signal for regulating their contributions, yet existing rollout-level methods largely treat disagreement as a scalar proxy for reliability. We analyze rollout-level uncertainty from an optimization perspective and identify two key limitations through theoretical and statistical analysis: the anisotropic gap, where probability dispersion fails to capture the geometric magnitude of semantic disagreement, and the calibration gap, where uncertainty is decoupled from reward informativeness and therefore cannot distinguish noisy variation from useful reward-differentiated variation. To address these gaps, we propose Geometric-aware Calibrated Policy Optimization (GCPO), which combines geometry-aware uncertainty measures with reward-based calibration to regulate policy updates according to both disagreement structure and learning informativeness. Across QA and mathematical reasoning benchmarks, GCPO delivers stronger GRPO-style optimization than entropy-based uncertainty baselines, while statistical and ablation analyses support the underlying design assumptions. Our results suggest a broader principle for uncertainty-aware policy optimization: uncertainty should characterize not only how much rollouts disagree, but also whether that disagreement provides useful learning signals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.