Rethinking On-Policy Distillation of Large Language Models: One Training Example
Abstract
On-policy distillation (OPD) trains on student rollouts with dense teacher supervision, but the role of training data remains unclear. We find that a single query supports hundreds of improving updates and recovers most of full-data OPD's gain across domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure state coverage: the fraction of the states full-data OPD visits that a query set's rollouts reach. One query reaches 71.5%, mostly within 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach 98.9% and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Rollouts quickly expose broad supervision early, while the student absorbs it increasingly slowly. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. These findings suggest that training queries matter through the teacher-supervised states their rollouts induce, while learning progress depends on how efficiently the student absorbs that supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.