MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
Abstract
The core challenge for streaming video generation is maintaining content consistency over long context, which poses high requirement for the memory design. Most existing solutions maintain the memory by compressing historical frames with predefined strategies. However, different to-generate video chunks should refer to different historical cues, which is hard to satisfy with fixed strategies. In this work, we propose MemFlow to address this problem. Specifically, before generating the coming chunk, we dynamically retrieve the most relevant historical frames as context with the text prompt of this chunk from the global memory bank. This design not only accurately sources the context needed to maintain visual consistency, but also ensures semantic coherence even as new events unfold or scenes transition. In addition, during generation, we only activate the most relevant tokens in memory for each query in the attention layers, which effectively guarantees the generation efficiency. After tuning the model on curated memory-intensive scenarios to unlock its ability to leverage distant historical cues, MemFlow achieves outstanding long-context consistency with minimal computation overhead (7.9% speed reduction compared with the memory-free baseline) and is designed to be compatible with any streaming video generation model with KV cache.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.