ReTaCo: Residual-Target Control for On-Policy Distillation
Abstract
On-policy distillation trains a student with teacher feedback on student-generated prefixes, but communicating or storing full-vocabulary teacher distributions at every generated token is costly. Entropy-aware on-policy distillation (EOPD) supplements reverse Kullback-Leibler (KL) divergence with forward supervision over the teacher's top- tokens to help recover plausible tokens that the student underestimates. Renormalizing the retained teacher probabilities assigns all target mass to the selected tokens and none to the omitted vocabulary. We prove that the loss continues to push the student's selected mass toward one even after the relative probabilities within the selected set match the teacher's. Consequently, matching the teacher is not a stationary point whenever the omitted vocabulary has positive teacher probability. This finding motivates ReTaCo (Residual-Target Control for On-Policy Distillation), which combines a single-sample estimator whose expectation equals the full-vocabulary reverse KL with a forward target that represents the top- tokens individually and groups all remaining tokens into one residual symbol. For teacher selected mass , the residual target is , while the relative probabilities of the selected tokens remain unchanged. Thus, preserves teacher mass, and larger values explicitly allocate more target mass to selected tokens. At a fixed prefix, we prove that the population objective has a unique optimum, whose selected mass lies between and and increases monotonically with . At , the forward term also retains non-vanishing recovery gradients for underestimated selected tokens. Numerical optimization confirms the predicted optimal mass and conditional distributions. Extensive experiments demonstrate the superiority of ReTaCo, outperforming EOPD on most mathematics and code benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.