acceptodds
Under review as a conference paper at ICLR 2027

KB-OPD: KL-Bridge On-Policy Distillation with Behavior-Corrected Prefix Guidance

Abstract

On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories, while generalized OPD (G-OPD) extends this framework through reward extrapolation. However, both depend on the prefixes visited by the student, which can limit access to productive reasoning trajectories early in training. Teacher guidance can improve these prefixes, but incorporating teacher-guided tokens into policy updates requires accounting for the mismatch between the sampling and student distributions. We propose KB-OPD, which combines constrained prefix sampling with reward extrapolation to learn from both guided prefixes and student-generated continuations. During a short warm-up phase, KB-OPD geometrically blends student and teacher distributions, selecting the largest teacher influence permitted by local and cumulative KL budgets. These constraints control both token-level and sequence-level deviation from the student, while a stopping rule hands generation back to the student for the remaining response. Clipped behavior importance weights and a weighted prefix–suffix objective incorporate bridge tokens into policy updates alongside student-generated tokens. After warm-up, sampling returns to pure student rollouts while retaining the G-OPD supervision signal. Experiments on mathematical reasoning and code generation show improved average performance over OPD and ExOPD across single-teacher, multi-teacher, and strong-to-weak distillation settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.