acceptodds
Under review as a conference paper at ICLR 2027

Distillation with Errors: Efficient On-Policy Distillation via Sparse Unbiased Gradient Estimation

Abstract

On-policy distillation (OPD) mitigates the distribution mismatch between training and inference by using teacher supervision on trajectories generated by the student. However, standard full-distribution OPD backpropagates dense gradients over the entire vocabulary, incurring substantial computational cost in the output-head backward pass. To reduce this cost, we introduce Distillation with Errors (DWE), an efficient OPD algorithm based on sparse unbiased gradient estimation. Our key insight is that distillation learns to match the teacher distribution over the course of training, so each step need not use the exact full-vocabulary gradient as long as the full-distribution gradient is preserved in expectation. Based on this insight, DWE exploits the zero-sum structure of the logit gradient to form two sampling distributions from its positive and negative parts. Sampling one coordinate from each distribution yields a paired sparse update whose expectation matches the original dense gradient. To reduce estimation error, DWE uses output-head geometry to guide the pairing. It then applies stratified sampling over pairs to reduce sampling variance. We theoretically establish the unbiasedness of the resulting estimator. Our experiments show that DWE reduces output-head backward time by up to 96.8% and total output-head time by up to 37.6%. Over 100 training updates, DWE reduces mean recorded training-step time by 2.31% while achieving average performance comparable to standard full-distribution reverse-KL distillation across 13 mathematics, code, and science benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.