VGSSC: Surface-Aware Voxel–Gaussian Representations for Monocular Semantic Scene Completion
Abstract
Monocular semantic scene completion seeks to infer a complete semantic occupancy field from a single RGB observation, requiring precise reconstruction of observed surfaces and reliable completion of unseen scene regions. This challenge becomes particularly pronounced in cluttered indoor scenes, where local geometry is complex and large portions of the scene remain unobserved. Existing methods typically represent scenes using either voxels or Gaussian primitives. Voxels preserve location-specific information throughout the volume but often produce spatially fragmented predictions, while Gaussians promote local geometric continuity yet may blur semantic boundaries when primitive supports overlap. We introduce VGSSC, a surface-aware Voxel–Gaussian framework that combines geometry-guided feature sharing with location-specific volumetric reasoning. Specifically, we introduce surface-relative distance tokens to encode the signed depth offset of each voxel from the predicted surface, providing explicit geometric cues that distinguish free space, near-surface regions, and occluded space. These surface-relative cues are incorporated into voxel features and also guide the initialization of surface-anchored Gaussian primitives. A Gaussian-assisted voxel decoder subsequently transfers locally aggregated geometric and semantic information from the Gaussian representation back to the dense volumetric, enabling volumetric completion beyond directly observed regions. Experiments on Occ-ScanNet and EmbodiedOcc-ScanNet demonstrate that VGSSC achieves state-of-the-art performance, including a 5.28% mIoU improvement over the previous method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.