MatchVR: Bridging the Reference Utilization Gap in Generative Video Restoration
Abstract
Generative video restoration can produce abundant details from severely degraded video, but its strong prior may hallucinate textures inconsistent with the true scene, especially when severe degradation removes reliable high-frequency evidence. Reference-guided restoration reduces this by anchoring reconstruction to real textures from high-quality reference frames, yet reference availability does not imply reference utilization: simply appending reference frames to a video diffusion transformer leaves much of the reference unused, even when reliable correspondences still exist. To address this gap, we present MatchVR, a one-step reference-guided generative video restoration framework that converts reference content into frame-specific matched guidance. Specifically, we introduce a degradation-robust feature matching strategy to estimate dense target-reference correspondences and their confidence. Additionally, we propose a spatiotemporal reference alignment module to warp and relocate the matched content to each target frame. We further introduce aligned-reference distillation to transfer full-reference teacher behavior. Experiments show that MatchVR consistently improves reference utilization and restoration quality across varying temporal distances, with particularly strong gains in fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.