KB-OPD: KL-Bridge On-Policy Distillation with Behavior-Corrected Prefix Guidance
Abstract
On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories, while generalized OPD (G-OPD) extends this framework through reward extrapolation. However, both depend on the prefixes visited by the student, which can limit access to productive reasoning trajectories early in training. Teacher guidance can improve these prefixes, but incorporating teacher-guided tokens into policy updates requires accounting for the mismatch between the sampling and student distributions. We propose KB-OPD, which combines constrained prefix sampling with reward extrapolation to learn from both guided prefixes and student-generated continuations. During a short warm-up phase, KB-OPD geometrically blends student and teacher distributions, selecting the largest teacher influence permitted by local and cumulative KL budgets. These constraints control both token-level and sequence-level deviation from the student, while a stopping rule hands generation back to the student for the remaining response. Clipped behavior importance weights and a weighted prefix–suffix objective incorporate bridge tokens into policy updates alongside student-generated tokens. After warm-up, sampling returns to pure student rollouts while retaining the G-OPD supervision signal. Experiments on mathematical reasoning and code generation show improved average performance over OPD and ExOPD across single-teacher, multi-teacher, and strong-to-weak distillation settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.