acceptodds
Under review as a conference paper at ICLR 2027

EFFICIENT ON-POLICY DISTILLATION OF LONG REASONING VIA OFF-POLICY PREFIX DISTILLATION

Abstract

On-policy distillation (OPD) trains a student on its own reasoning with dense teacher feedback. However, the teacher must supervise long reasoning traces that it would not have written itself, which makes the student learn slowly. We observe that the accuracy of the student greatly increases when it continues from the beginning of the reasoning generated by the teachers, and vice-versa the performance of the teacher is degraded when it has to start from a partial reasoning traces written by the student. This large effect of the reasoning prefix motivates our proposed OPD variant, **Prefix-Supervised On-Policy Distillation (PS-OPD)**. At each training step, the student continues a reasoning prefix generated offline by the teacher, and it is trained by supervised fine-tuning on the prefix and OPD on its own continuation. In this way, the student will generate higher quality answers and receive more informative supervision from the teacher, as demonstrated by faster convergence and better final performance of PS-OPD compared to OPD. For example, on AIME24, AIME25 and AIME26 with a 7B math teacher and 1.5B student, PS-OPD reaches in training steps the accuracy OPD attains after steps, ends points higher, and needs about fewer GPU-hours for training, including the generation of teacher traces. The gains carry over to a teacher fine-tuned by reinforcement learning from the student itself: PS-OPD leads OPD over the first ten steps and ends points higher. PS-OPD is also more robust to the choice of distillation loss: across four forward- and reverse-KL losses, its final accuracy spans and points with the two teachers, against and for OPD. It also curbs OPD's response-length explosion: with the RL-finetuned teacher under the main loss, up to of OPD's training rollouts reach the length cap, against at most of PS-OPD's.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.