Decoupled On-Policy Distillation for LLM Reasoning
Abstract
On-policy distillation (OPD) commonly relies on forward or reverse KL divergence (FKL/RKL) to align the student with the teacher. However, FKL has a particular limitation in OPD: its per-token loss and gradient concentrate on a small set of candidate tokens on which the teacher probability far exceeds the student probability . We call this high-disparity set of candidates the Mismatched Region, where exceeds by one to two orders of magnitude, and this concentration makes OPD prone to response-length inflation, unreasonable truncation, and accuracy collapse. RKL, in contrast, avoids this domination through its -self-weighting but can under-cover teacher modes through its mode-seeking behavior. To address these issues, we propose a simple but effective optimization objective, Decoupled-KL Distillation (DKLD), a loss that decomposes the KL into two contributors, the Matched Region and the Mismatched Region. In detail, we simply use forward KL in the Matched Region, and Reverse KL with -self-weighting takes responsibility within the Mismatched Region. Furthermore, the student rollouts required by OPD simultaneously provide the group-relative advantages for GRPO, but their gradients can conflict. To reduce this interference, we propose an advantage-weighted distillation module that balances these two objectives and combines OPD with GRPO in a single framework. Through comprehensive experiments, across mathematical reasoning and code generation tasks, DKLD reaches state-of-the-art performance in all configurations we ran and drastically mitigates the length-inflation collapse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.