Memento: Reconstruct to Remember for Consistent Long Video Generation
Abstract
Long-form video generation requires recurring subjects to remain consistent across shots, viewpoints, and scene transitions. Memory-conditioned autoregressive methods improve scalability through shot-by-shot generation. However, they primarily optimize memory conditioning or selection mechanisms, without explicitly supervising the utilization of identity-critical evidence in historical memory, allowing subject appearance to drift over time. We propose Memento, a subject-reconstruction-guided framework that treats subject preservation as an explicit identity grounding problem, introducing memory-based subject reconstruction to encourage the model to preserve and retrieve identity-critical appearance cues. Memento jointly trains next-shot generation and subject reconstruction, recovering target appearances from historical memory and subject descriptions. A dual-query memory mechanism separates these historical cues: story-conditioned queries retrieve long-range subject evidence, while shot-conditioned queries select short-range contextual references for coherent continuation. A subject-aware cinematic data pipeline aligns story, shot, and reconstruction captions through consistent subject descriptions, providing unambiguous supervision for identity grounding. Experiments demonstrate state-of-the-art subject consistency, particularly across scene transitions, together with strong cross-shot coherence and competitive visual quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.