Coefficient-Calibrated Memory Retrieval via Sparse Coding for Long Video Generation
Abstract
Autoregressive video generation models use fixed context windows to make long rollouts tractable, but progressively lose early information as earlier frames are evicted from the context. Existing retrieval methods retrieve the top- historical latent entries from a memory bank by similarity, improving subject and background consistency. However, independent top- selection may retrieve near-duplicate historical latents whose accumulated evidence causes *memory over-anchoring*, suppressing motion evolution and ultimately leading to motion collapse. We propose a training-free memory interface to address this failure, namely the Coefficient-calibrated Atom REtrieval of MEMory (CARE-MEM). We map historical latent entries in the memory bank to memory atoms that form an online dictionary. Our approach adopts constrained sparse coding to jointly select a fixed-budget set of memory atoms and applies the resulting coefficients to reweight the attention evidence of the corresponding memory entries before normalization. Across three autoregressive backbones and -, -, and -second generation under VBench-Long, the proposed method achieves the best or second-best performance in of metric–setting comparisons. Compared with the top- retrieval baseline, our method improves Dynamic Degree by points on average while maintaining comparable Motion Smoothness, with additional gains in quality and consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.