Harnessing Video Diffusion Model as Prior and Supervision for Video Restoration
Abstract
Adversarial transfer has proven an effective way of exploiting pretrained diffusion priors for image restoration: the restoration model is initialized from a diffusion prior and optimized adversarially against a discriminator that is itself a strong pretrained image model, which yields single step restoration at high perceptual quality. Extending this recipe to video is not straightforward. A pretrained video diffusion model can initialize the restoration network, but pretrained image discriminators act frame by frame, and a comparably capable pretrained video discriminator is not readily available to judge spatial appearance and temporal dynamics jointly. We, therefore, use distribution matching with the score of a pretrained video generation model to supervise temporal dynamics without requiring a video discriminator. This also avoids VSR-specific diffusion pretraining, reducing training costs and complexity. For spatial supervision, we apply image adversarial training to one randomly sampled frame per clip, whose gradient largely supports the full-frame image adversarial update. Extensive experiments demonstrate that the proposed HYPVR achieves state-of-the-art spatial quality and temporal consistency with a single forward pass, making effective use of the prior for efficient video restoration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.