acceptodds
Under review as a conference paper at ICLR 2027

VoxScene: Voxel Representation for 3D Scene Arrangement

Abstract

Current data-driven layout generation techniques typically encode objects using bounding-box proxies or implicit features. While computationally efficient, such representations abstract away volumetric occupancy, obscuring object boundaries and weakening direct geometric constraints on spatial arrangement. To address these limitations, we introduce VoxScene, which shifts the 3D scene arrangement paradigm to an explicit voxel representations, enabling direct modeling of fine-grained spatial relations. Our method factors the generation process into two stages: First, we generate box-level anchors to specify the coarse layout. Second, an object-centric diffusion network sequentially synthesizes instance-labeled volumetric occupancies conditioned on these anchors. Specifically, each voxel is assigned to at most one instance, preventing inter-instance overlap at the voxel level. These generated voxels further serve as geometric queries for asset retrieval. Extensive experiments on 3D-FRONT demonstrate that our method outperforms existing methods in physical plausibility while maintaining visual quality. Furthermore, results on M3D-Shelf show that VoxScene can generate high-quality arrangement for fine-grained objects under complex structural constraints, where previous methods struggle.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.