Response-Disagreement Preconditioning for Causal Video Distillation
Abstract
Distilling a bidirectional video diffusion model into a few-step causal generator requires transferring text-conditioned behavior across fundamentally different temporal factorizations. The teacher denoises with full-sequence context, whereas the student makes sequential, prefix-conditioned predictions. In asymmetric distribution matching distillation (DMD), the causal student is updated using the discrepancy between an online estimate of its rollout score and a classifier-free-guided teacher score. This discrepancy provides the base correction for distribution matching, but does not explicitly characterize that correction relative to differences in how the teacher and student score fields respond to text conditioning. We introduce Null-Relative Response-Preconditioned DMD (NRP-DMD) to exploit this response disagreement within the existing DMD update. At the same noised student rollout and diffusion timestep, we evaluate both score models under the target text and a null condition. For each model, the null-to-text change defines a text-induced score response, and the disagreement between the teacher and student responses defines a sample-dependent axis in score space. NRP-DMD uses this axis to precondition the existing guided DMD update, increasing the gain of its response-aligned component while leaving the orthogonal component unchanged. NRP-DMD modifies only the distillation procedure and requires neither architectural changes nor additional inference-time computation. Across two causal distillation baselines, NRP-DMD consistently improves generation quality over both 5- and 30-second horizons while maintaining strong text alignment and temporal coherence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.