acceptodds
Under review as a conference paper at ICLR 2027

Structured Motion-Instance Graphs and Persistent Memory for Long-Horizon Video Generation

Abstract

Generating long videos in successive chunks requires preserving instance identity, motion, and interactions across chunk boundaries. We propose a framework that couples a Motion-Instance Graph (MIG) with persistent Instance Memory (IM) for controllable long-horizon video generation. MIG organizes subjects, static regions, trajectories, interactions, and event states into a typed graph, which is converted into frame-level motion and spatial conditions. IM retrieves historical instance states before each chunk and updates them from generated observations, allowing subsequent chunks to reuse appearance, trajectory, and interaction information. Lightweight MI-Adapter and LoRA modules inject graph conditions and memory tokens into a frozen video diffusion backbone. First-frame anchoring and frame-aware corruption further support cross-chunk consistency. We evaluate the framework on 160 conditions with three random seeds, conduct incremental ablations, and test generation at 240, 720, and 1,440 frames. On CogVideoX-2B, normalized average displacement error decreases from 0.2108 for the trajectory-control baseline to 0.1850, while background consistency increases from 0.618 for the long-video baseline to 0.750. Evaluations on Wan2.1-1.3B and VideoCrafter2 also show improvements over their paired baselines. The framework provides explicit instance-level control and persistent state for long-video generation without updating the backbone parameters.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.