Teaching a Moving Student: Rethinking the Curriculum of On-Policy Distillation
Abstract
In on-policy distillation (OPD), the student not only learns from the teacher but also determines which states receive future supervision. As the student evolves, training moves to new response prefixes while teacher–student disagreement can remain on earlier ones. Under matched trajectory and optimization budgets, neither more training queries nor more frequent rollout resampling is uniformly beneficial. Fresh current-policy rollouts outperform initial-policy rollouts at shorter response budgets, but the ranking reverses at longer budgets. Fixed-prefix measurements show that earlier prefixes become less likely under the student while teacher–student disagreement on them persists. Replaying initial-policy rollouts after current-policy training improves accuracy, whereas replaying fixed recent rollouts does not reproduce the gain. We therefore propose R-OPD, a gradient-triggered curriculum that lets the student learn on current-policy states before adaptively introducing replay, without prescribing a fixed transition iteration. Once mean gradient changes fall within minibatch-level variation, training switches to initial-policy rollouts from the next iteration through the remaining training budget. Across eight mathematics benchmarks, R-OPD improves average accuracy over continued current-policy sampling by 2.25 and 4.44 percentage points at 16K and 32K for a 0.6B student across three training runs, and by 2.81 and 6.15 points for a 1.7B student. With a 30B-A3B teacher, R-OPD also improves 8B accuracy by 4.05 points at 32K. Fixed two-stage replay also improves accuracy. At 32K, R-OPD exceeds a fixed schedule of 40 current-policy updates followed by 20 replay updates by 1.39/1.66/1.85 points for 0.6B/1.7B/8B, averaged over three training-data orders. At the same 32K generation cap, R-OPD also produces longer responses, suggesting that well-timed revisits help the student use more of its reasoning capacity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.