acceptodds
Under review as a conference paper at ICLR 2027

Beyond Mode-Covering and Mode-Seeking: A Closer Look at KL Divergence in On-Policy Distillation

Abstract

On-policy distillation (OPD) is a de facto recipe to enhancing the reasoning capability of student models by matching the teacher's distribution over self-generated trajectories, where the choice of divergence is a critical design decision. Empirical guidance relies on geometric intuitions — framing forward KL as *mode-covering* and reverse KL as *mode-seeking* — offering little insight into how these objectives actually steer optimization. Moving beyond this dichotomy, we shift our focus to optimization dynamics. Specifically, we analyze the logit-level gradient fields that different divergences produce during training. And we show that no -divergence guarantees that every logit update moves probability toward the teacher due to the Softmax coupling effect. Instead, effective optimization relies on coordinate-level sign alignment with the student-teacher probability gap. Under this framework, we prove that forward KL is the unique sign-aligned -divergence, while reverse KL actively pushes over-allocated tokens in the wrong direction — a structural pathology that convex combinations of the two cannot resolve. To settle this, we introduce m-RKL, a minimal correction of RKL's gradient field that restores sign alignment while retaining RKL's beneficial tail damping. Extensive experiments on Qwen3 family models demonstrate that m-RKL outperforms competitive KL-based objectives on a suite of math reasoning benchmarks without degrading the general capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.