TOKEN UTILITY FOR ON-POLICY DISTILLATION VIA LOCAL TRUST-REGION OPTIMIZATION
Abstract
Token-selection methods for on-policy distillation (OPD) prioritize student-visited positions using diverse signals, including uncertainty, teacher–student disagreement, and mismatch direction. We instead derive token importance from the distillation objective itself: how much can each position's loss decrease under an equal local KL budget on the student's predictive distribution? The answer is the utility , where is the logit gradient and is the categorical Fisher matrix. For reverse-KL it equals the student-weighted variance of the student–teacher log-probability ratio, and for forward-KL it equals the Pearson divergence, so the appropriate utility follows from the choice of objective. On shared training checkpoints, several independently motivated criteria, including student entropy, KL disagreement, and TIP, concentrate their importance on nearly the same positions as the derived utility at a 50% selection fraction, while diverging at stricter fractions and for direction-specific criteria. We instantiate the utility as Fisher Utility Weighting (FUW), a continuous, mean-normalized token weighting. Across seven mathematical reasoning benchmarks and three Qwen3 student–teacher pairs, FUW improves over matched uniform OPD under both reverse-KL (1.35–4.0 macro points) and forward-KL (1.08–2.71 macro points), and is competitive with existing selective methods using a single closed-form score derived from the training objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.