acceptodds
Under review as a conference paper at ICLR 2027

CoVoGen: Coarse Voxels for Compositional 3D Scene Generation

Abstract

Compositional 3D scene generation from sparse observations requires complete geometry and coherent spatial arrangement. We introduce CoVoGen, a scene-first framework that extracts coarse object voxels from a completed shared scene to guide component refinement. Given posed RGB-D views with instance masks, a scene predictor completes coarse occupancy in world coordinates. A learned membership classifier extracts each target's coarse voxels from this shared prediction. Object structure refinement combines these voxels with neighboring occupancy and aligned observations to predict finer occupancy. A frozen pretrained geometry generator converts the refined occupancies into meshes in the shared scene. On four benchmarks comprising 1,129 scenes and 6,880 target objects, CoVoGen achieves the highest F@5 and Mean F for objects and scenes in our main comparison using the same input views, improving scene F@5 by up to 4.57 percentage points over the best baseline. These results support compositional scene generation through scene-derived coarse structure and object refinement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.