What Does On-Policy Distillation Optimize?
Abstract
On-policy distillation uses teacher feedback on student-generated prefixes, so an update changes both the student and the distribution of its later supervision. Yet practical methods often replace sequence gradients with token-local, bounded, support-restricted, or future-truncated updates. This raises a basic question: *what objective do these updates represent?* We answer this question at three levels: First, every continuous coordinate-local reward admits an explicit separable potential; we characterize its proper and Bregman subclasses. Next, universal autoregressive composition within the separable Bregman class forces the logarithmic generator, exposing a tradeoff between bounded rewards and sequence consistency. Finally, under student occupancy, an exact future-credit decomposition yields an intrinsic objective-fidelity coefficient, a finite-step descent guarantee, and circulation tests for integrability. To validate these results, exact finite trees verify these identities and separate composition from update fidelity. Complementing these exact checks, Qwen3.5 measurements on GSM8K, ARC-Challenge, and HellaSwag, using 0.8B-to-4B and 4B-to-9B model pairs, reveal horizon-dependent transitions from descent to near orthogonality or ascent. Matched training further shows that cached-prefix teacher fit and task utility can rank the same updates differently. Together, these results identify the objective each surrogate represents, the sequence terms it omits, and when its local update preserves progress.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.