PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet each rollout also reveals how the student’s response unfolds and whether it succeeds, hindsight that standard OPSD does not use to form the teacher. We introduce Privileged Adaptation from Student Trajectories (PAST), which treats each completed student trajectory as additional privileged information for the OPSD teacher while leaving the student’s distillation prefixes unchanged. PAST preserves the student’s next-token distribution on correct trajectories and adapts the teacher on failed trajectories toward verified success while keeping it close to the student. We characterize what a teacher informed by the trajectory can transfer to a prefix-only student. Forward-KL distillation projects the teacher distributions to their conditional arithmetic mean given the prefix. This projection separates variation across trajectories that remains privileged from the mean policy shift available to the student. For correct trajectories, the unclipped population objective also has the frozen student as an ideal distributional fixed point. Across three mathematical reasoning benchmarks, PAST improves the Avg@12 macro average over Vanilla OPSD by 5.6 percentage points. A 2×2 factorial study shows gains from complete trajectories and teacher adaptation, while trajectory removal and shuffling confirm that the adapted teacher uses the matching hindsight context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.