Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
Abstract
Knowledge distillation has become central to LLM post-training, yet its design space remains poorly understood, especially alongside reinforcement learning (RL). We show that the prevailing paradigms, off-policy distillation and on-policy distillation (OPD), implicitly **couple two orthogonal choices: prefix source and token-level KL direction**. This coupling follows from decomposing sequence-level KL over autoregressive response distributions: *forward KL pairs teacher prefixes with token-level forward KL, and reverse KL pairs student prefixes with token-level reverse KL*. We argue that this coupling is not intrinsic: decoupling the two axes yields four valid objectives. We establish gradient-level identities showing that forward KL gives SFT-style cross-entropy matching with teacher soft targets, whereas reverse KL gives an RL-style policy-gradient objective with a dense teacher-student log-ratio reward, connecting the four objectives to off-policy SFT, DAgger-style on-policy SFT, offline-RL-style distillation, and OPD. We conduct an extensive controlled study on math reasoning, evaluating the four objectives both as standalone distillation methods and as initializations for subsequent RL. The results reveal three tradeoffs: *KL direction induces an accuracy–entropy tradeoff, prefix source induces a quality–compute tradeoff, and training length induces an accuracy–stability tradeoff*. Motivated by these findings, we propose **KL mixing** and an **entropy-gated length curriculum**. KL mixing shows that long-sequence distillation requires substantial forward-KL weight to prevent entropy collapse and length inflation without sacrificing accuracy. The entropy-gated length curriculum improves Avg@k and Pass@k by **3.6** and up to **5.8** points, and reduces average response length by roughly relative to fixed long-horizon training. Our framework and methods guide the design of reasoning distillation objectives that balance accuracy, diversity, compute, and RL behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.