acceptodds
Under review as a conference paper at ICLR 2027

A Unified KL– Geometry of Approximate Policy Mirror Descent

Abstract

Regression-based reinforcement learning for language models approximates policy mirror descent (PMD). The exact KL-regularized update requires a log-partition function that is difficult to estimate from a few rollouts. Practical methods modify this baseline, as in the two-temperature formulation of \APO and the mean baseline of Kimi K2. We study how these choices change the policy-improvement objective. For regression relative to the rollout policy, we characterize a mixed KL– objective with the same population optimum. The induced strength is zero at matched baseline and update temperatures, increases continuously with baseline temperature, and reaches its maximum at the mean baseline. We bound this strength using reward variation and characterize its dependence on the two temperatures and rollout pass rate for binary rewards. The bounds identify penalty strengths that cannot be reached through baseline temperature alone. We then optimize the mixed objective directly, making the strength an explicit parameter. Experiments with Qwen2.5 models on MATH and DAPO-Math-17k show that baseline choice and explicit penalty strength affect training stability. On MATH, positive coefficients sustain training in settings where KL-only updates fail across all three model sizes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.