acceptodds
Under review as a conference paper at ICLR 2027

Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion

Abstract

Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness, as exemplified by the Self Forcing training paradigm. However, existing autoregressive video diffusion models still incur substantial attention and memory overhead because they retain largely redundant key–value (KV) caches for historical frames, limiting their scalability. In this paper, we tackle this challenge by introducing KV cache compression into autoregressive video diffusion. We observe that attention heads in mainstream AR diffusion models exhibit markedly distinct attention patterns and functional roles that remain stable across samples and denoising steps. Building on our empirical study of head-wise functional specialization, we divide the attention heads into two categories: static heads, which focus on transitions across autoregressive chunks and intra-frame fidelity, and dynamic heads, which govern inter-frame motion and consistency. We then propose Forcing-KV, a hybrid KV cache compression strategy that performs structured static pruning for static heads and dynamic pruning based on segment-wise similarity for dynamic heads. While maintaining output quality, Forcing-KV compresses 53%-73% of the KV cache, achieves 29 FPS on a single NVIDIA H200 GPU with 30% peak cache memory reduction, and delivers up to 1.35x and 1.50x speedups on LongLive and Self Forcing at 480P resolution. The efficiency gains further scale with resolution and sliding window size. Code and demo videos are provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.