OPD-Cookbook: Practical Recipes for Reliable and Effective On-Policy Distillation
Abstract
On-Policy Distillation (OPD) combines on-policy training with dense token-level supervision from a stronger teacher, making it an increasingly important component of modern model post-training. Despite its simple formulation, making OPD effective in practice is not always straightforward. To address this challenge, we propose an OPD-Cookbook consisting of three complementary algorithmic recipes for reliable and effective OPD. The Cookbook starts from the mechanism of OPD and organizes its analysis around two fundamental aspects of the distillation process: the supervision signal provided by the teacher and the objective used to align the teacher and student. From the perspective of teacher supervision, we first introduce Teacher Adaptation on OPD Data. When the teacher retains improvement headroom on the data used for OPD, further adapting it with reinforcement learning can strengthen the supervision subsequently provided to the student. We further introduce Teacher Recovery Training on Student States, motivated by the observation that a teacher capable of solving a problem on its own trajectory may still provide unreliable guidance when conditioned on imperfect student prefixes. We therefore explicitly train the teacher to recover from states visited by the student. From the perspective of the alignment objective, we propose a Decoding-aware KL Objective that aligns the token distributions actually used during decoding rather than only the raw model distributions, better matching the training objective with generation behavior. Across our experiments, these three recipes provide complementary improvements over standard OPD and together offer a practical set of principles for building more effective OPD pipelines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.