ResGVC: Residual Coding for High-fidelity Generative Video Compression
Abstract
Generative video compression (GVC) methods usually employ decoder-side diffusion models to denoise reconstructed pictures based on generative priors. Besides, most of these methods are applied in latent space rather than directly in pixel domain. In this paper, we propose ResGVC, a residual coding framework for GVC that guides denoising toward original pictures by transmitting residual signals. ResGVC first encodes the source pictures within a learned video codec and then decomposes the distortion between the codec reconstruction and the source pictures into a sequence of residual segments, where each residual segment is acquired at each denoising step. Residual segments are then compressed using a novel residual codec with a diffusion transformer as entropy model. Unlike prior methods, ResGVC introduces explicit source constraints throughout denoising rather than relying on decoder-side generation alone. The residual codec also enables more efficient context modeling capability based on the powerful generative priors of diffusion transformer. Moreover, operating directly in pixel space avoids the irreversible information loss caused by latent representation bottlenecks. Experimental results show that, compared with the previous state-of-the-art method, ResGVC achieves 63.7% and 83.8% BD-rate reductions in terms of LPIPS and PSNR, respectively, while also delivering superior visual quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.