acceptodds
Under review as a conference paper at ICLR 2027

GenMem: Learning Ultra-Compact Video Memory by Generative Supervision

Abstract

Recent video generation models can produce long videos, but maintaining long-term consistency remains challenging because conditioning on extensive visual history requires prohibitive memory and computation. We introduce GenMem, a compact visual memory for long-context video generation. GenMem can compress 4 video frames of 480p resolution into only 32 tokens while retaining high-fidelity visual information. Our memory combines flexible Q-Former-based token compression with a hierarchical representation that prioritizes important semantics, appearance, and dynamics within a small token budget. To train our memory encoder, we find that supervision from the generation loss enables more aggressive compression while better preserving semantically relevant information than using vanilla reconstruction loss. This compact design allows video generators to efficiently condition on several minutes of visual history. Experiments show that GenMem enables substantially more scalable long-context generation while preserving the consistency of scenes and objects observed from distant video history.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.