acceptodds
Under review as a conference paper at ICLR 2027

Harnessing Video Diffusion Model as Prior and Supervision for Video Restoration

Abstract

Adversarial transfer has proven an effective way of exploiting pretrained diffusion priors for image restoration: the restoration model is initialized from a diffusion prior and optimized adversarially against a discriminator that is itself a strong pretrained image model, which yields single step restoration at high perceptual quality. Extending this recipe to video is not straightforward. A pretrained video diffusion model can initialize the restoration network, but pretrained image discriminators act frame by frame, and a comparably capable pretrained video discriminator is not readily available to judge spatial appearance and temporal dynamics jointly. We, therefore, use distribution matching with the score of a pretrained video generation model to supervise temporal dynamics without requiring a video discriminator. This also avoids VSR-specific diffusion pretraining, reducing training costs and complexity. For spatial supervision, we apply image adversarial training to one randomly sampled frame per clip, whose gradient largely supports the full-frame image adversarial update. Extensive experiments demonstrate that the proposed HYPVR achieves state-of-the-art spatial quality and temporal consistency with a single forward pass, making effective use of the prior for efficient video restoration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.