SURGE: Sparse Update Routing via Gated Entropy for Efficient Active Distillation
Abstract
On-policy distillation (OPD) is widely used to transfer capabilities from large teacher language models to smaller student models. However, supervising every token incurs substantial computational cost without distinguishing where guidance is truly needed and can induce unnecessary updates that disrupt existing student capabilities. To address these limitations, we introduce , an active distillation framework that provides selective control over student adaptation and knowledge transfer by using two separate gating mechanisms for updates. A token gate applies teacher supervision only at high-entropy positions, while a step gate forces updates to remain inside low-scoring generation steps. Additionally, a decoupled KL objective is used to apply updates selectively on such active tokens while retaining regularization toward a frozen reference student over the entire sequence. Across the evaluated reasoning, scientific knowledge, and hate speech tasks, SURGE achieves the highest reported accuracy among the compared baselines. When distilling Gemma4-31B into Gemma4-E4B, SURGE improves over the base student by a relative on average and achieves a mean recovery of of the initial teacher–student accuracy gap. Relative to OPD, estimated distillation FLOPs fall by on average, and the share of that gap SURGE closes is larger. Average response length falls by , while the mean format-failure rate falls by relative to OPD. Finally, we assess the generalization of our method by distilling Qwen3.5-122B-A10B into Qwen3.5-2B on these benchmarks, also noting improvements across accuracy and training FLOPs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.