VAEM: Give Video Avatars an Evolving Mind
Abstract
Recent efforts have explored endowing video avatars with active intelligence to generate multi-step video sequences toward high-level goals. However, existing methods fail to consolidate past failures into persistent knowledge, frequently repeating similar errors in planning and execution. To address this, we present VAEM (Video Avatar Evolving Mind), a framework that distills video-generation histories into a persistent planner context comprising behavioral principles in an evolving system prompt and reusable skills in a structured library. VAEM operates via a Plan-Generate-Critique-Update (PGCU) cycle, orchestrating a scene-aware planner, a frozen video generator, specialized vision-language critics, and an update agent. During context evolution, the critics localize semantic failures, enabling the update agent to convert feedback into actionable rules and skills for subsequent tasks. At inference, the accumulated context is frozen and reused directly for planning without online updates. Extensive experiments on half-body and full-body manipulation tasks demonstrate that the six-round CCE setting reduces mean semantic failure counts from 9.26 to 7.83 and 8.85 to 7.65, respectively, while enhancing aggregate visual quality. Furthermore, on the L-IVA benchmark using a shared generation backend, VAEM achieves state-of-the-art perceptible motion amplitude (PAS: 28.0) and motion smoothness (MSS: 87.7) alongside competitive visual quality. Finally, a real-time prototype validates VAEM's compatibility with interactive rendering. These results demonstrate that persistent context evolution offers an effective paradigm for cross-task experience reuse without fine-tuning the underlying video generator. Project page: https://anonymous.4open.science/w/VAEM-04EB.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.