acceptodds
Under review as a conference paper at ICLR 2027

Bootstrapped On-Policy Distillation

Abstract

On-policy distillation (OPD) improves language models by distilling teacher predictions on student-generated trajectories, while on-policy self-distillation (OPSD) removes the need for an external teacher by using a reference-conditioned copy of the model as the teacher. Despite its great early improvements, we find that OPSD does not remain consistently effective under continued training and can eventually suffer substantial performance degradation. We study this training behavior and link the degradation to a growing mismatch between the evolving reference-free student and the frozen privileged teacher, together with increasing student uncertainty on its evolving on-policy states. This growing mismatch motivates us to first learn from privileged reference supervision and then transfer the acquired knowledge to a fresh student. We introduce Bootstrapped On-Policy Distillation (BOPD), a two-stage framework that first uses OPSD to internalize task-relevant information from reference-conditioned supervision into an intermediate model, and then distills this reference-free model into a fresh copy of the initial model. Experiments across 4B and 8B models show that BOPD improves mathematical reasoning beyond both the initial models and the constructed OPSD teachers, while also outperforming the constructed teachers on code generation. These results suggest that privileged reference supervision can be more effectively exploited by first transferring its task-relevant signal into a task-adapted model and then distilling from that model without reference conditioning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.