MemID: Memorizing Identities in Streaming Narrative Video Generation
Abstract
Memory is central to streaming autoregressive (AR) video generation, where later events of a multi-prompt narrative must stay consistent with content that has fallen outside the bounded context window. Existing memory methods mainly focus on what to remember, yet they either overwrite early identity anchors through continuous memory updates or rely on pretrained LLMs/VLMs as external retrievers in inference. More importantly, they overlook where to render each identity: even with correct identity references, overlapping attention between different identities causes ID fusion, while fragmented attention within the same identity leads to ID duplication. We propose MemID, an identity-aware streaming AR video generation framework that addresses both aspects. For what to remember, MemID introduces an ID-preserving Memory that compresses each prompt-conditioned event into representative frames, freezes it after the event ends, and retrieves only prompt-relevant entries for current generation. For where to render, MemID is trained with ID-aware Attention Losses that separate different identities while encouraging compact and temporally coherent attention for each identity. Extensive experiments demonstrate that MemID substantially improves long-range identity consistency and maintains accurate instance counts in multi-character narratives, while preserving competitive generation quality. Code and models will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.