When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better
Abstract
On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask: Is on-policy sampling always beneficial for distilling arbitrary teacher–student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher–student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. It also beats other OPD variants on low-overlap pairs. We further find that the advantage of Semi-OPD strongly depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned. Our study suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering all reasoning stages and exposing the student to more teacher-preferred tokens. Our work motivates rethinking when to use OPD based on teacher–student overlap and adopting Semi-OPD as an efficient and strong alternative when appropriate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.