What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation
Abstract
On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can sometimes attenuate useful supervision, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student's reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and restores additional supervision for prefix-dependent tokens. Trajectory dropout consistently improves performance across teacher–student model pairs of different scales and six mathematical reasoning benchmarks. It can also be directly integrated into existing OPD variants, including TRD and FastOPD, further improving their performance. These results demonstrate that trajectory dropout provides a simple and general mechanism for strengthening token-level supervision across model scales and OPD objectives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.