acceptodds
Under review as a conference paper at ICLR 2027

World-δ: Continually Integrating Observations for Incremental 3D Scene Generation

Abstract

Existing 3D scene generation methods can synthesize visually compelling scenes from fixed inputs, yet struggle to continually incorporate new observations to coherently expand and revise an existing 3D world. Existing iterative approaches often fuse real observations and generated content into a single deterministic representation without explicitly modeling their provenance or uncertainty. Consequently, early generation errors can accumulate and become difficult to correct with subsequent observations. To address these limitations, we propose World-, an incremental 3D scene generation framework with uncertainty-aware persistent world modeling and delta-guided diffusion for continual scene expansion and correction. Our framework consists of three components: 1) Visual Evidence Encoding integrates geometric and semantic cues from incoming images into unified world tokens. 2) Observation-Calibrated World Modeling maintains a persistent latent world state that tracks geometry, appearance, provenance, observation confidence, and generative uncertainty, distinguishing observation-supported content from provisional completion hypotheses. It compares new evidence with the existing state to infer selective world deltas, enabling the revision of conflicting hypotheses while preserving well-supported content. 3) Incremental World Diffusion Generation derives update masks from these deltas and restricts diffusion to regions requiring generation or revision, while freezing unchanged regions. Only affected 3D chunks are decoded and assembled with the existing scene representation. Together, these components enable the world to expand with accumulating observations and allow uncertain completions to remain revisable, while reducing redundant generation and preserving scene consistency. Extensive experiments demonstrate the effectiveness of World- in incremental 3D scene generation, observation-driven correction, scene consistency, and update efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.