TSD: Few-Step Video Diffusion with Temporally Decoupled Distribution Matching Distillation
Abstract
Reliable video generation is central to interactive applications and digital content creation. Although distribution matching distillation (DMD) enables few-step video generation, balancing visual fidelity with coherent motion remains challenging during prolonged training, and temporal variation may even collapse. We present Temporal Subspace Distillation (TSD), a general supervision framework designed to make better use of teacher knowledge during continued training. Our key idea is to decompose video latents into a temporal mean replicated across frames and zero-mean deviations. These components lie in orthogonal subspaces, enabling separate supervision of persistent visual content and temporal variation. Building on this decomposition, we introduce an identity anchor to transfer time-constant content from the teacher and one-sided deviation floors to penalize insufficient temporal variation in the student. We further extract motion saliency from the same teacher predictions to guide where variation is protected without prescribing a specific motion trajectory. Experiments show that TSD improves video quality over baseline methods, producing videos with high visual fidelity and coherent, plausible motion. TSD retains the student architecture and inference procedure and adds no inference cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.