AOSR: Asynchronous On-Policy Speculative Rollout for Efficient Distillation of Autoregressive Image Generators
Abstract
On-policy distillation (OPD) improves compact autoregressive image generators by training them on image-token sequences generated by the current student. However, producing these sequences autoregressively requires repeated sequential forwards, making student rollout a major training bottleneck. We explore speculative multi-token rollout to reduce this cost while maintaining distillation performance, but find that a batch-synchronous speculative implementation can even reduce efficiency: different sequences obtain different amounts of verified progress, while synchronous execution prevents faster sequences from advancing independently. Based on this observation, we propose AOSR, an Asynchronous On-policy Speculative Rollout framework that preserves sequence-specific verified progress and adaptively removes completed sequences. We further introduce Cross-Step Speculative Refill (CSR), which reuses completed batch positions to pre-generate sequences for the next rollout and treats the prefetched tokens as speculative candidates. After the student update, these tokens are re-verified with the updated student, resampling mismatched tokens when necessary, thereby reusing computation across OPD steps while maintaining comparable generation quality. Experiments on LlamaGen and Janus-Pro with GKD and VarKD show that AOSR substantially reduces rollout and end-to-end training time, while adding CSR provides further acceleration with comparable generation quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.