TraceWorld: Trajectory-Indexed Memory for Spatially Consistent Video World Model
Abstract
Camera-controlled video world models aim to simulate coherent visual environments across changing viewpoints. When the camera revisits a previously observed region, maintaining this coherence requires scene information that may no longer be available in recent temporal context. We introduce TraceWorld, a video world model with view-adaptive key-value (KV) memory allocation for camera-controlled generation. To incorporate camera geometry while preserving the pretrained visual pathway, TraceWorld adds a decoupled camera-aware attention branch. Pose-affinity retrieval then selects historical latent frames and reuses their cached KV states across both attention branches. A learned allocation policy divides a fixed context budget between recent and retrieved KV states, conditioned on inter-chunk camera-pose changes and the pose affinity between target views and stored history. Together, these components balance local temporal continuity with scene recall without constructing an explicit 3D scene memory. To evaluate scene consistency under camera revisits, we introduce the Scene Recall Benchmark (SRBench), comprising 200 camera-return cases across eight motion categories. On WorldScore, TraceWorld achieves the highest object control and 3D consistency scores among the compared methods. On SRBench, it achieves the highest revisit consistency and margin while remaining competitive in stability and camera following.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.