Eviction Is Not Forgetting: Read Long, Write Local in Streaming Video Diffusion
Abstract
Block-autoregressive video diffusion models generate minute-long videos one short latent block at a time. Each block is denoised while attending to a key/value (K/V) cache of earlier blocks; the finished block then passes through the network once more to compute the K/V that the cache stores, and a fixed window evicts the oldest block. We find that eviction does not remove an old block’s influence. Because this writing pass attends to the current cache, every new entry stores a copy of older entries, and the copy stays after its source is evicted: for fixed video, a cache difference in an -layer model with a -block window can last up to writes instead of . Writing each block alone removes the copies but also loses how the video entered the block. We therefore read long and write local. ARROW keeps the full cache for denoising and writes each entry from the clean latents of the previous and current blocks, without old K/V, storing only the current block’s K/V. It needs no training and keeps one extra latent block. On Self Forcing (Wan2.1), MAGI-1 and HunyuanVideo-1.5, ARROW raises VBench-Long quality by 3.6–6.2 points and LoCoT2V character consistency by 10.6–29.4 points. Replay and writer ablations trace these gains to removing the copies and adding the boundary frames, and reader masking shows that long-range reading must be kept.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.