Spatial‑geometric Self‑supervised Learning for LiDAR Panoptic Segmentation
Abstract
Recent advances in self‑supervised learning (SSL) for LiDAR point clouds have significantly improved 3D scene understanding without requiring human annotations. However, existing approaches still face challenges in capturing fine‑grained geometric structures from highly sparse outdoor LiDAR scenes, especially under limited annotation settings. Learning robust geometric representations is essential for LiDAR panoptic segmentation, which requires both accurate semantic recognition and precise instance localization. In this work, we introduce SpaGeo3D, a spatial‑geometric self‑supervised framework for LiDAR representation learning and panoptic segmentation. SpaGeo3D learns geometric‑aware representations from unlabeled LiDAR scans through masked voxel reconstruction within visible local neighborhoods, enabling the model to capture spatial structures and surface geometry while avoiding unnecessary computation on large empty regions. To further exploit complementary spatial information, we propose a Multi‑scale Polar Feature Fusion module that effectively integrates sparse 3D geometric features with polar‑domain contextual representations. For instance‑level prediction, an offset‑density refinement mechanism jointly leverages geometric displacement and BEV density cues to improve instance center localization and object‑level grouping. Extensive experiments on the SemanticKITTI and nuScenes datasets demonstrate that SpaGeo3D consistently outperforms existing approaches, achieving average improvements of +7.2% PQ and +12.2% mIoU in panoptic segmentation, while exhibiting strong generalization ability in large‑scale outdoor LiDAR scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.