Temporal Re-Distillation: A Post-Distillation Stabilizer for One-Step Video Super-Resolution
Abstract
One-step video super-resolution (VSR) models distilled via Distribution Matching Distillation (DMD) can preserve the frame-wise perceptual quality of multi-step diffusion teachers, yet often exhibit pronounced temporal artifacts, including flickering and texture drift. To understand this discrepancy, we trace temporal refinement along the teacher's denoising trajectory and observe that fine-grained cross-frame consistency is refined predominantly at low noise levels, a regime largely bypassed by one-step distillation. Motivated by this observation, we propose TRD-VSR, a temporal refinement framework that re-perturbs student predictions to low-noise states and distills the teacher's local denoising corrections in this regime. We further introduce Relational Structure Alignment (RSA) to explicitly transfer cross-frame relational structure, complementing the local teacher corrections with an explicit constraint on temporal correspondence. TRD-VSR is used only during training and therefore preserves single-step inference with no additional test-time cost. Experiments on three benchmarks show that TRD-VSR improves both temporal consistency and perceptual quality, achieving the best tLPIPS across all datasets, the best tOF and tLP on two datasets, and the best CLIPIQA and DOVER across all datasets, yielding a favorable perception–consistency trade-off over existing one-step generative VSR methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.