LadderOPD: Logit-Free On-Policy Distillation from the Teacher’s Top-k Order
Abstract
On-policy distillation (OPD) trains a small model on its own samples while a teacher scores every token with its full next-token distribution. Many strong teachers are served through interfaces that return only sampled text or a short list of top candidates, so this distribution is often unavailable. We study how much of it OPD actually needs. We first introduce GroundKL, a measurement that takes an open teacher’s logits as ground truth and tests how well each signal a closed interface could return recovers the exact per-token OPD signal. The signal is concentrated at high-entropy forking tokens, and there the order of the teacher’s top-k candidates recovers it nearly as well as sixteen samples from the teacher. Guided by this finding, we propose LadderOPD, which trains on the teacher’s ranked top-k token IDs alone. Across five teacher–student settings in two model families, LadderOPD is statistically equivalent to OPD with top-k probabilities or full logits within three GSM8K points in all ten comparisons, and it matches them on code as well. Shuffling the order at a growing share of forking tokens degrades the student step by step. For verifiable reasoning with a shared tokenizer, a teacher’s ranked top-k list is therefore enough for on-policy distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.