WMem: Relevance, Reliability, and Timing for Long-Horizon Video World Memory
Abstract
Video world models can generate realistic short sequences. However, they often lose scene identity as observations leave the local context and generation errors accumulate. We argue that effective long-horizon memory is not only a storage problem, since its utility depends on which observations are retrieved, how reliable the memory is, and when it is read during denoising. We introduce WMem, a memory framework that collectively addresses these three factors through Role-factorized Retrieval, Diffused Memory Conditioning, and Denoising-Stage Routing. WMem retrieves stable anchors, target-view revisits, and complementary views, and combines them with role and target-relative geometry. During training, our Diffused Memory Conditioning trains the model to use memories with query-dependent fidelity without expensive autoregressive rollouts. Denoising-Stage Routing allocates memory roles across denoising stages, reducing unnecessary computation. With the same backbone and memory budget as the baseline, our WMem reduces MSE by 36.1%, LPIPS by 35.6%, and FID by 82.0%, while improving PSNR by 3.68 dB over Minecraft 500-frame rollouts. It also reduces inference time by 23.5% and peak memory of the GPU by 36.9%. On 300-frame forward-and-return RealEstate10K trajectories, WMem also improves long-horizon scene consistency. Our results show that long-horizon world modeling benefits from jointly controlling memory relevance, reliability, and timing. Code and models will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.