acceptodds
Under review as a conference paper at ICLR 2027

GENERATIVE VIDEO COMPRESSION FOR SPATIAL DETAIL AND TEMPORAL CONSISTENCY

Abstract

Generative video compression relies on visual priors to restore detail at low bitrate, while maintaining temporal consistency across frames remains challenging. Frame-wise synthesis with image generative priors can produce inconsistent textures even with temporal conditioning, resulting in flickering artifacts. In this paper, we present GVC-ST, a generative video compression framework for rich spatial detail and stable temporal consistency. GVC-ST divides the video sequence into a reference chain and non-reference frames with distinct responsibilities. The reference chain propagates recovered detail and maintains temporal consistency across the sequence, where each frame is reconstructed by one-step diffusion with forward attention to the preceding reconstructed reference frame. For non-reference frames, each frame is reconstructed from its neighboring references and can optionally be replaced by frame interpolation, achieving high detail without affecting the reference chain. Experimental results on three benchmarks show that GVC-ST achieves the best FID among all compared generative codecs while delivering substantially less temporal flickering than prior generative approaches, as measured by and .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.