When KL Regularization Misfires in Group Policy Optimization
Abstract
Why does removing reference-policy KL regularization sometimes improve group policy optimization? We argue that this may stem from misaligned KL regularization and imbalanced aggregation, and analyze seven potential failure modes: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which coordinates KL regularization with group-relative reward updates through the construction and joint integration of KL coefficients. Mathematical reasoning experiments and ablations support the effectiveness of ZCPO in our settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.