Flash On-Policy Distillation
Abstract
Full-vocabulary KL provides on-policy distillation (OPD) with supervision over every next-token alternative, rather than a teacher-top-k truncation or a sampled-token estimate. Making this objective efficient is a framework-level challenge: dense distribution materialization strains memory and communication, while complete-response barriers delay updates and serialize teacher work. We introduce Flash-OPD, an OPD training framework that makes full-vocabulary KL a native training path through three coordinated designs. Partial on-policy rollouts retain unfinished prefixes as context and optimize only fresh suffixes generated by the current student snapshot. Teacher overlap evaluates each ready suffix while other requests continue generating, preserving one synchronized student update per round. A tiled KL loss kernel reconstructs teacher logits from hidden states and a frozen output head, using online normalization and analytic backward recomputation instead of materializing the complete token–vocabulary matrix. Together, these designs preserve full-vocabulary supervision and token-level behavior-policy matching while reducing update latency and loss workspace, with FSDP and vocabulary-parallel Megatron support. In a representative wall-clock comparison, Flash-OPD reaches reverse KL 0.01 in 4.2 hours versus 8.5 hours for full-rollout OPD, a 2.02× time-to-target speedup, with comparable final divergence. Efficient OPD need not obtain its systems benefits by truncating the vocabulary supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.