KAVA: Keyframe-Anchored Video Adapters for Extreme Video Compression
Abstract
Video codecs, classical and learned, transmit whatever their decoder cannot predict from the frames it has already decoded, and at low rates they drop fine details. A pretrained video-diffusion model has learned much of how natural video looks and moves, and if the same copy sits at both ends of the channel, it can predict much of what a codec would otherwise transmit. We present KAVA, a codec that specializes one such frozen model, Wan2.1, to each target video. A standard codec sends sparse, heavily compressed keyframes, and the model generates every frame from them, the keyframe positions included. Because the model alone would produce plausible motion rather than the target's, we fit a rank-one adapter to each video on the layers where tokens enter and leave the transformer and on its text and time conditioning, leaving every transformer block frozen. The bitstream is these keyframes and this adapter alone. On full-length UVG sequences, KAVA delivers state-of-the-art perceptual performance from to bits per pixel against traditional, learned and prior generative codecs. We see specializing such a shared model to each video, so that the bitstream carries only what the model cannot predict, as a key step toward extreme video compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.