Learnable Compressed Memory for Autoregressive Video Generation
Abstract
Autoregressive video diffusion transformers generate videos chunk-by-chunk with causal attention over a sliding-window KV cache. While the window bounds computation and memory, it discards early context and can lead to long-horizon drift. Frame-sink variants retain a few anchor frames, but often introduce temporal discontinuities and abrupt regressions toward anchor-like appearances. In this paper, we present , a learnable memory system that eliminates the need for frame sinks while preserving access to a compact, evolving summary of long-range history under a fixed token budget. trains a KV Compress Adapter end-to-end with the base model and uses perceiver-style cross-attention with learnable latent queries to select and summarize KV tokens, compressing KV tokens to represent 7 more history while maintaining attention fidelity. Compressed memories are organized in a fixed-capacity memory bank with exponential temporal binning and consolidated through content-aware eviction and merging, yielding a diverse and temporally balanced summary without hard boundaries. Experiments on long-form video generation show that achieves long video generation with better visual quality, mitigates drift, avoids sink-induced artifacts, and preserves dynamic motion with smooth camera movement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.