Reverse-KL Policy Distillation: Statistical Guarantees and Limits
Abstract
everse Kullback–Leibler (KL) distillation has shown empirical promise for train- ing language models, but its statistical guarantees and limits remain poorly under- stood. In particular, the effects of two design choices remain unclear: Does access to the teacher’s full action distribution provide an advantage over observing only the probability it assigns to the sampled action? Must training trajectories come from the student (online distillation), or can teacher-generated trajectories (offline distillation) suffice? Our work takes a first step toward understanding the theory underlying these questions. We establish matching upper and lower bounds for online reverse-KL distillation when the teacher reveals its full action distribution, and show that substantially more data may be required when the teacher reveals only the probability of the student-generated action. In contrast, offline reverse-KL distillation can attain the same optimal estimation rate whether the teacher reveals its full distribution or only its probability of the teacher-generated action. The distinction is that teacher- generated actions themselves reveal the teacher’s preferences, while queries at student-selected actions can leave those preferences unresolved. This highlights an additional statistical challenge for online distillation compared to offline distillation. These statistical guarantees do not follow simply from optimizing standard dis- tillation objectives: our lower bounds show that common online objectives can fail when candidates are evaluated only on past student-generated trajectories. We develop trajectory-loss and importance-ratio truncation techniques together with policy-space learning to achieve optimal online distillation guarantees.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.