EvoMA: Evolving Entity Memory Agent for Consistent Long Video Generation
Abstract
Consistent long video generation requires preserving entity identity while tracking evolving states. Existing planning-based methods typically predefine entity-state trajectories and visual references, which can diverge from the actual content of generated clips and continuing user inputs. We introduce EvoMA, an Evolving Entity Memory Agent that treats memory management as an online decision-making process. EvoMA maintains a structured memory that separates persistent identities from evolving states and records state histories for characters, scenes, and objects using textual descriptions and visual references. A vision-language agent updates and queries this memory through an observe–think–act–verify loop, dynamically performing add, update, noop, and retrieve operations based on newly generated clips. Memory changes are verified against visual evidence before being committed, while retrieved entity-state references guide subsequent generation. We further train the memory agent with reinforcement learning to improve both memory updates and state retrieval. Experiments show that EvoMA improves identity consistency and state alignment in long and interactive video generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.