ChunkOPD: Efficient Strict On-Policy Distillation via Chunk-Wise Execution
Abstract
On-policy distillation (OPD) trains student language models with teacher feedback on their own generations, aligning supervision with the states they visit. For long reasoning trajectories, however, conventional strict OPD execution incurs teacher-scoring waits and substantial computation and activation-memory costs during student training. Asynchronous execution improves utilization but introduces policy staleness that may degrade distillation quality, motivating a more efficient schedule that preserves strict on-policy property. We present ChunkOPD, which exploits prefix causality and student parameter consistency between rollout and training to reorganize both teacher scoring and student training around token chunks. Streaming Chunk Scoring scores committed prefixes before trajectories finish, overlapping teacher computation with rollout to shorten the scoring tail. Reverse Chunk Training reuses rollout key–value (KV) states to avoid an additional cache-population pass and recomputes chunks in reverse order to propagate gradients through the cached context while keeping activations local to each chunk. Across the evaluated workloads under matched physical GPU budgets, ChunkOPD achieves 1.11–1.98 speedups over synchronous OPD. In an end-to-end case study, ChunkOPD reduces time costs by 41% and achieves almost identical quality relative to synchronous OPD, whereas asynchronous OPD attains lower final accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.