acceptodds
Under review as a conference paper at ICLR 2027

The Design Space of On-Policy Distillation: A Unified Framework and Empirical Analysis

Abstract

On-policy distillation (OPD) transfers reasoning capabilities from a stronger teacher to a student by training on student-generated responses. Yet it remains unclear which parts of OPD are essential for effective learning and which can be simplified. We introduce a unified training objective that organizes representative OPD methods along four design axes: rollout source, trajectory weighting, token-position weighting, and local supervision loss. Guided by this framework, we conduct a systematic empirical study of over 30 OPD configurations on mathematical reasoning tasks. We find that teacher feedback remains effective on student-generated trajectories. In our setting, training only on failed trajectories matches the accuracy of using all trajectories. At the token level, a random 100-token window remains competitive with full-token supervision, while SelecTKD, which emphasizes teacher–student agreement, achieves the highest accuracy among the tested token-weighting methods. Preserving teacher probability information remains important: adaptive vocabulary allocation retains more teacher probability mass at the same target average budget, whereas replacing probabilities with ranks or top- membership lowers accuracy. Skew KL improves accuracy using a mixture of teacher and student probabilities, while hard clipping limits extreme learning signals. Finally, we combine these choices into a single OPD recipe and evaluate it on three larger teacher–student model pairs, improving accuracy by up to points and reducing total training time by . Together, our framework and findings clarify where OPD can be simplified and where richer teacher feedback remains important.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.