Where Should Suppressed Probability Go? One-Step Lookahead KL Routing for On-Policy Distillation
Abstract
On-policy distillation (OPD) supervises a student on its own trajectories, but its sampled-token form leaves the recipients of negative credit implicit. When a sampled token is suppressed, the compensating logit update follows the student's current probabilities rather than explicitly selecting recipients using teacher evidence, even among alternatives that both models rank highly. We formulate this limitation as recipient routing and introduce Lookahead KL Routing (LKR). LKR first defines a dense routing operator. It combines current teacher support with agreement between the student and teacher after appending each candidate, then solves a constrained reverse-KL projection to reallocate a bounded portion of the update within this shared candidate set. Its unique water-filling target defines a routing correction that preserves the sampled token's local logit gradient. Dense routing would otherwise require successor evaluation at every eligible event. LKR makes training practical by sampling a small subset of events and adapting the projection budget to keep each routing update bounded. Across six public mathematical reasoning benchmarks, LKR improves step-matched OPD by 2.23 and 4.54 percentage points in Mean@8 and Cons@8, respectively. The gains persist across three seeds on a 1,000-problem held-out set with a 1.5B student and a 7B teacher. The pattern holds with a 7B student and a 32B teacher.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.