VOGVR: A Native Vision-Only Generative Video Restoration Model
Abstract
Most existing generative video super-resolution (VSR) or video restoration (VR) methods are built by adapting text-to-image (T2I) or text-to-video (T2V) diffusion models pretrained on large-scale text-visual data. However, these models often suffer from issues of temporal consistency (e.g., T2I-based methods) or detail synthesis (e.g., T2V-based methods). In addition, the synthesis-oriented nature of the T2I/T2V backbones, even finetuned on low-quality (LQ) and high-quality (HQ) paired data, can make it difficult to keep generated details tightly constrained by the observed input. We investigate whether a generative VR model trained purely on visual data can surpass these T2I and T2V-based adaptations, and propose **VOGVR**, a native **V**ision-**O**nly **G**enerative **V**ideo **R**estoration model. We first train a bidirectional VAE to capture temporal context, reducing latent ambiguity, and improving fine-detail reconstruction. Then, we extract visual features and motion cues from LQ videos using pretrained vision encoders (e.g., DINOv2) and optical flow networks, and use them as visually grounded conditions to guide the training of the DiT, anchoring the generation to the LQ input for detail-rich and temporally coherent restoration. A high-quality dataset with more than 90M images and 3M video clips is built to train our VOGVR model. Extensive experiments demonstrate that VOGVR outperforms state-of-the-art generative VSR/VR models, effectively alleviating the trade-off between temporal coherence and spatial details while keeping the restoration grounded in the input video, all with one-step inference. Moreover, VOGVR requires less than 1/30 of the training cost of representative T2I-based VR methods and less than 1/400 of T2V-based ones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.