FReD: Frequency-Resolved Distillation for Video Generation
Abstract
Few-step distillation accelerates video generation by reducing the number of denoising evaluations needed at inference. Distribution-matching approaches primarily achieve this by aligning teacher and student distributions in latent space, but such alignment does not explicitly assign teacher guidance to individual spatial frequency bands. To investigate this issue, we analyze generated and real videos in the spatial frequency domain and observe distinct discrepancies in their low- and high-frequency components. These discrepancies motivate frequency-resolved distillation with dedicated supervision for low- and high-frequency predictions. Specifically, we decompose the outputs of a frozen teacher model and a trainable fake score model into two frequency bands while retaining the complete input context, and the difference within each band supervises the corresponding component of the same native student prediction, extracted using a non-learnable filter. We further introduce overlap-aware loss weighting to regulate supervision in the transition region where the low- and high-frequency components overlap. Our method adds band-residual losses to the training code and incurs little training overhead, while preserving the original sampler and leaving the number of denoising steps and inference cost unchanged. Experiments across text-to-video, image-to-video, and autoregressive video generation tasks demonstrate improved generation quality and closer spectral agreement with real videos.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.