World-T : World-Test-Time Training as Persistent Memory for Real-Time Video World Models
Abstract
Recent advances in video world models have enabled plausible interactive environments, yet long-horizon generation places a fundamental demand on visual memory: the previously observed world must remain consistent as the model continues to generate. Existing autoregressive approaches often implicitly rely on the generative model itself to retain such information, making persistent scene memory difficult to maintain beyond the active context. We introduce World-T, a camera-controlled video world model with a test-time training memory that consolidates past observations into dedicated model parameters. To robustly update this memory, we propose a recursive least-squares rule that incrementally absorbs novel visual features via online pose-keyed memory updates. Therefore, our TTT memory enables consistent generation by reading out previously observed scene features via camera pose from the fast-weight state, while maintaining constant storage and per-step computation regardless of rollout length. To facilitate real-time interactive generation with this robust memory, we further distill the autoregressive model into a few-step generator via dual-teacher distribution matching. Experiments show that World-T achieves precise camera control and faithful scene revisits while sustaining 20.6 FPS at 1280 × 704 on a single NVIDIA B200 GPU.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.