Mitigating Blending Drift: Adaptive Blended Policy Distillation for LLM Reasoning
Abstract
Large language model (LLM) knowledge distillation (KD) with reinforcement learning (RL) has recently shown promising results in enhancing the reasoning capabilities of student models by transferring knowledge from stronger teachers. Existing approaches predominantly fall into two pure policy paradigms. Off-policy distillation uses teacher-generated rollouts for high-quality supervision, but may exceed the student's learning capacity and induce severe train–test mismatch. On-policy distillation preserves learnability by using student-generated rollouts, but suffers from prefix failure, where erroneous student prefixes hinder effective teacher supervision. Policy blending offers a natural compromise by combining teacher and student policies for rollout generation. However, we identify a previously overlooked failure mode, termed blending drift, in which the blended policy assigns probability mass to regions unsupported by either the teacher or the student. Through theoretical analysis, we provide a theoretical guarantee that reducing the sampling temperature effectively mitigates this drift. Motivated by this finding, we introduce Adaptive Blended Policy Distillation (ABPD), a novel and simple sharpen-then-blend strategy that generates high-quality rollouts while remaining within the student's on-policy support. Extensive experiments on mathematical reasoning and code generation benchmarks demonstrate that ABPD consistently achieves the best overall performance among competing methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.