acceptodds
Under review as a conference paper at ICLR 2027

When KL Regularization Misfires in Group Policy Optimization

Abstract

Why does removing reference-policy KL regularization sometimes improve group policy optimization? We argue that this may stem from misaligned KL regularization and imbalanced aggregation, and analyze seven potential failure modes: residual KL updates after reward clipping, after gradient cancellation, and in groups with identical rewards; KL growth with response length and an imbalance in its relative contribution; KL concentration on a small number of tokens; and sampling noise when k1 is incorporated into rewards. We propose Zero-Sum Calibrated Policy Optimization (ZCPO), which coordinates KL regularization with group-relative reward updates through the construction and joint integration of KL coefficients. Mathematical reasoning experiments and ablations support the effectiveness of ZCPO in our settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.