WorldSculpt: Generating Compositional Worlds from Grounded Videos
Abstract
Generating large-scale cluttered scenes as individually editable object meshes is challenging. Existing compositional methods largely target simple scenes. Our key insight is that, with appropriate design, a single-object generative prior can scale to large-scale, densely cluttered scenes without scene-level training. We introduce , which realizes this idea through canonical-space multi-view conditioning, training-time augmentation, and object-wise scene composition. Finetuned entirely on single-object data, it generates hundreds of mutually occluding objects as complete, spatially aligned meshes in a shared world frame. We further introduce a photorealistic benchmark of densely cluttered scenes with per-object annotations. Evaluations demonstrate superior geometric accuracy and observation alignment over prior approaches. We also demonstrate broader applicability by converting 3DGS worlds from Marble into compositional scenes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.