acceptodds
Under review as a conference paper at ICLR 2027

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning for agent

Abstract

On-policy distillation (OPD) aligns a student model with a teacher’s logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially less data. However, standard token-level OPD can provide only fragmented corrections along an erroneous student trajectory and cannot unfold a complete and correct repair path. Motivated by this limitation, we propose Step-Level On-Policy Distillation (SOPD), which combines the long-horizon correction of supervised fine-tuning (SFT) with the on-policy advantage of OPD to provide step-level supervision over complete student-generated trajectories. We show that, at different limits of step length, SOPD reduces to SFT or approximates OPD. Compared with SFT, the teacher responses in SOPD are conditioned on student trajectories and therefore align more closely with student-visited states; compared with OPD, SOPD provides longer-horizon corrections rather than fragmented token-level guidance. On ALFWorld, the Qwen2.5-7B student trained with SOPD improves Seen and Unseen success rates over SFT by 36.42 and 38.58 percentage points, respectively, and over Vanilla OPD by 15.34 and 15.17 points, while reducing action rounds. We hope this work offers a new perspective for future research on distillation methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.