ViableOPD: Guiding On-Policy Distillation with Prefix Viability
Abstract
On-policy distillation (OPD) transfers teacher capabilities by providing dense, token-level teacher supervision on student-generated trajectories. This paradigm implicitly assumes that the entire trajectory can benefit from teacher supervision. However, once a student-generated prefix deviates from the task objective, teacher supervision is conditioned on an erroneous context, making distillation at such states ineffective or even harmful. To characterize when teacher supervision remains effective, we introduce prefix viability, defined as the probability that the original task objective can still be achieved from a student-generated prefix. We estimate prefix viability by normalizing the teacher’s success probability conditioned on the prefix by its success probability from the original input. Based on this criterion, we propose Viability-Guided On-Policy Distillation (ViableOPD), which identifies viable rollouts for teacher supervision, adjusts token-level supervision strength within selected rollouts, and adapts the rollout horizon according to the fraction of viable rollouts. Across multiple benchmarks and teacher–student pairs, ViableOPD consistently outperforms strong OPD baselines while reducing training time by 33.6% on average relative to standard OPD. These results suggest that effective OPD requires prefixes to be not only visited by the student, but also viable for continued supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.