L3G-World: Real-Time Infinite World Generation with Latent 3D Gaussian Memory
Abstract
Interactive world models generate visual content conditioned on user actions, yet maintaining spatial consistency over long-horizon exploration in real time remains challenging. Existing approaches rely on either frame retrieval or explicit 3D memories, both of which scale poorly as memory grows with the generated stream. Moreover, explicit 3D memories condition latent video generators by re-encoding RGB renderings, which adds computation and propagates rendering artifacts. In this paper, we introduce a Latent 3D Gaussian Memory (L3G) that combines the efficiency of feed-forward reconstruction with a compact, rapidly renderable scene representation. Building on this memory, we propose \ours, in which a memory updater and a video diffusion transformer alternate autoregressively to generate spatially consistent worlds. The memory updater spawns Gaussians in newly observed regions and refines existing ones in explored regions, so that the memory grows with scene extent rather than stream length. Each Gaussian further stores a VAE latent feature, allowing the memory to be rendered directly into the DiT's latent space without VAE re-encoding. Experiments show that \ours achieves state-of-the-art generation quality and long-horizon spatial consistency, with up to fewer 3D primitives, faster memory update, and faster memory rendering than prior 3D memory-based methods, enabling real-time world generation at 9 FPS.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.