Theoretical Analysis of On-Policy Distillation: Generalization and Pass@k
Abstract
On-policy distillation trains a student using teacher feedback on its own generated responses. We provide a theoretical analysis of its generalization and its Pass@k performance. First, we establish a uniform finite-sample upper bound relating distillation error to empirical training loss, sample size, student description length, and the shift in prefix distributions between training and evaluation. A complementary lower bound quantifies the unavoidable error from unobserved prefixes and recovers the same dependence on sample size and distribution shift in a worst-case construction. Second, we establish upper and lower bounds on the student-teacher Pass@k gap that match in order when KL error is small. Finally, we theoretically show that an exact KL proximal update can reduce distillation error and improve Pass@1 while lowering Pass@k. We identify a mechanism for this trade-off: teacher preferences can suppress correct responses on harder prompts, while gains on easier prompts raise Pass@1 but contribute little at larger sampling budgets. Synthetic experiments verify our theory, and experiments on large language models support the practical relevance of our analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.