acceptodds
Under review as a conference paper at ICLR 2027

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

Abstract

On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask: Is on-policy sampling always beneficial for distilling arbitrary teacher–student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher–student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. It also beats other OPD variants on low-overlap pairs. We further find that the advantage of Semi-OPD strongly depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned. Our study suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering all reasoning stages and exposing the student to more teacher-preferred tokens. Our work motivates rethinking when to use OPD based on teacher–student overlap and adopting Semi-OPD as an efficient and strong alternative when appropriate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.