NoiseFlow: Warp-Guided Video Generation with Distilled Temporal Priors
Abstract
Temporal coherence in video diffusion models is often imposed using motion correspondences estimated in video space, implicitly assuming that pixel-space motion defines an appropriate transport in latent noise space. We challenge this assumption and introduce NoiseFlow, a framework that learns the temporal noise structure directly from a pretrained video diffusion model. NoiseFlow distills the denoiser-induced local geometry into diffusion-timestep-conditioned autoregressive warp fields that transport Gaussian noise across video frames. We then train a student diffusion model on the resulting family of warped noise priors. During inference, the warp is inferred from the evolving denoising trajectory, and the student provides warped-trajectory guidance to the base model, enabling video generation without a reference video or externally estimated optical flow. Experiments show that NoiseFlow reduces FVD by up to 18%, improves SSIM by up to 43%, and PSNR by up to 19%, and improves several metrics of VBench and VBench2 compared to the unguided base model. These results suggest that denoiser-native temporal priors provide a useful alternative to prescribing video-space correspondences in generative noise space.\footnote{Models and code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.