Sharp Horizon Scaling of Gradient Variance in On-Policy Distillation
Abstract
On-policy distillation (OPD) trains a student on self-generated sequences using positionwise teacher feedback. Temporal credit rules pair feedback with scores differently, but their tight horizon dependence remains unclear. Using a score–feedback matrix and conditional score centering, we show that past links affect variance without changing the mean, while future links can carry sequence credit. Under bounded feedback and scores, arbitrary prefix dependence, and shared parameters, we establish tight worst-case total variance of for token-local and for full-return and reward-to-go (RTG) in unnormalized rollout updates. One shared-parameter OPD construction attains all three rates at a fixed nonzero teacher–student gap, tightening the previous quartic full-return bound. Our control analysis characterizes removable variance and approximation costs; even optimal time-only controls retain cubic worst-case sequence variance. Full-parameter Qwen and Llama measurements show nearly linear token-local variance and faster sequence variance growth, revealing different length sensitivities. In controlled GPT-2 training with matched objectives and rollout budgets, RTG improves optimization, reducing normalized final reverse KL from 0.73 to 0.35 at . This advantage persists without gradient clipping.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.