WorldNoise: Making Video Diffusion Models See the 3D World through Geometric Noise
Abstract
Recent advances in camera-controlled video diffusion models have enabled interactive world modeling, where an initial noise is progressively transformed into 3D scenes that can be explored along arbitrary camera trajectories. Yet, surprisingly, the source of these 3D worlds—randomly sampled diffusion noise—is agnostic to scene geometry. This raises a fundamental question: can the diffusion noise itself serve as a structured representation of the 3D world? In this paper, we introduce WORLDNOISE, a novel framework that brings 3D structure directly into the noise space of video diffusion models. Our key insight is to reinterpret diffusion noise not merely as unstructured randomness, but as a latent representation of the world that can be structured and transformed according to 3D geometry. Specifically, WORLDNOISE first constructs a 3D noise space over the observable world region jointly covered by the camera frusta along the trajectory and distributes randomness throughout this space. Then, it projects the world-space randomness into view-specific 2D noise through a carefully designed whitening and covariancecompensation process that recovers the native spatial noise prior expected by the video diffusion model. Different views therefore become different observations of the same underlying noise world, rather than being generated from unrelated noise. The resulting WORLDNOISE can be used through the existing diffusion noise interface without modifying the denoiser architecture. We further extend WORLDNOISE to long-horizon exploration by allowing the noise world to persist and grow as the camera moves through the scene. Extensive experiments demonstrate the effectiveness of WORLDNOISE, revealing a previously overlooked role of diffusion noise as a latent carrier of 3D structure and opening a new direction for 3D-aware generative world models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.