MotionMem: Future Trajectory Forecasting as Dynamic Memory for Video Generation
Abstract
Long-horizon world generation requires models to both preserve past observations and reason about how the world evolves over time. Existing memory mechanisms are effective for recalling static scenes, but are less suited to dynamic entities whose future states must be inferred before historical information can be reused. We propose MotionMem, a unified plan-retrieve-generate framework that explicitly couples future-state prediction with memory retrieval. MotionMem jointly models RGB videos and instance-level track maps, predicting future trajectories with only a single denoising step as a lightweight motion plan, which is then used to retrieve and align relevant historical appearance for subsequent RGB generation. Trajectory-Guided Dynamic Memory uses the predicted motion to retrieve and place historical object appearance alongside camera-aligned scene memory. We further distill MotionMem into a causal few-step streaming generator for efficient interactive generation. Experiments demonstrate improved long-term consistency in both scene recall and dynamic evolution, while preserving these advantages under real-time streaming generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.