VoRTeC: Taming Foundation Flow for one-step Real Time Video Compression
Abstract
Ultra-low bitrate video compression still faces critical challenges: traditional neural video compres- sion inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor temporal consistency. To address these issues, we propose VoRTeC, a Video Compression framework built upon a foundational flow model (Wan2.1). By compactly encoding latent video representations, predicting the positions of compressed represen- tations along flow trajectories, and integrating multi-scale priors, VoRTeC enables the compressor to harness generative video flow priors effectively. Without accessing the parameters or gradients of flow matching networks, our framework achieves one-step decoding and reconstructions with high perceptual fidelity. Meanwhile, we maintain consistency across frame groups via tail-frame reuse and prior caching. Extensive experiments demonstrate that our method reduces bit consumption by 58% compared to prior diffusion-based approaches, with decoding speed boosted by 3 to 197 times: VoRTeC achieves a decoding speed of 13 FPS at 720p and 32 FPS at 480p.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.