acceptodds
Under review as a conference paper at ICLR 2027

EncoreMem: Re-encounter Verification of Region Memory for Training-Free Long Video Generation

Abstract

Autoregressive video models generate a long video one chunk at a time, conditioning each chunk on a cache of what has already been generated. Keeping the earliest frames in this cache helps limit degradation. These frames hold only what appeared at the start, so the rest of the cache must decide which later content to keep and for how long. Existing methods choose what to keep by its current or predicted relevance, or by learned representations. We instead keep stored content only if later frames, which also follow the current prompt, return to it. We introduce EncoreMem, which stores newly appearing parts of frames after they persist over several chunks and keeps each part provisionally. Whether a part returns is detected with the features of the generator, and when memory fills, the part that returned least often in recent chunks is replaced. The rule needs no training, additional model, retrieval at read time, or cache rebuild when the prompt changes. In minute-long video generation with both multi and single prompts, EncoreMem achieves competitive overall performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.