acceptodds
Under review as a conference paper at ICLR 2027

Where Should Suppressed Probability Go? One-Step Lookahead KL Routing for On-Policy Distillation

Abstract

On-policy distillation (OPD) supervises a student on its own trajectories, but its sampled-token form leaves the recipients of negative credit implicit. When a sampled token is suppressed, the compensating logit update follows the student's current probabilities rather than explicitly selecting recipients using teacher evidence, even among alternatives that both models rank highly. We formulate this limitation as recipient routing and introduce Lookahead KL Routing (LKR). LKR first defines a dense routing operator. It combines current teacher support with agreement between the student and teacher after appending each candidate, then solves a constrained reverse-KL projection to reallocate a bounded portion of the update within this shared candidate set. Its unique water-filling target defines a routing correction that preserves the sampled token's local logit gradient. Dense routing would otherwise require successor evaluation at every eligible event. LKR makes training practical by sampling a small subset of events and adapting the projection budget to keep each routing update bounded. Across six public mathematical reasoning benchmarks, LKR improves step-matched OPD by 2.23 and 4.54 percentage points in Mean@8 and Cons@8, respectively. The gains persist across three seeds on a 1,000-problem held-out set with a 1.5B student and a 7B teacher. The pattern holds with a 7B student and a 32B teacher.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.