Memory-Gated Distillation: Mitigating Self-Generated Memory Drift in Autoregressive Diffusion World Model Distillation
Abstract
Autoregressive diffusion world models enable interactive long-horizon video generation, but their sampling cost is incurred at every generated chunk, making low-latency deployment difficult. Few-step distillation from a bidirectional teacher to an autoregressive student is a promising remedy, but rollout distillation becomes unstable over long horizons. We identify self-generated memory drift as a key cause of this instability: early student errors are written into memory and reused as conditioning context, degrading subsequent distillation. We introduce Memory-Gated Distillation (MGD), which addresses this error-accumulation problem through selective memory supervision. MGD uses a tailored chunk-level reliability verifier to decide whether each student-generated chunk should be written into memory, thereby preserving reliable rollout prefixes while filtering corrupted ones. As a result, MGD retains the benefits of on-policy student rollouts while suppressing supervision from corrupted memory contexts, with no additional inference-time cost. Experiments on short- and long-horizon generation demonstrate improved few-step generation quality and a stronger quality-latency tradeoff than prior distillation baselines. Particularly in the challenging 253-frame rollout setting, 4-step MGD nearly matches the 50-step bidirectional teacher on aesthetic quality (0.564 vs. 0.569) and PSNR (13.81 vs. 13.84) with less than 1/10 of sampling time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.