Self-Supervised Representation Learning from Lidar Images
Abstract
In real-world applications, such as autonomous driving, the inherent complexity of the environment poses high requirements on the quality of the representations learned by perception models. While semantic understanding is important, it is usually not present natively in LiDAR data. To solve this, for the first time, we investigate self-supervised learning for LiDAR range-view representation with the hypothesis that distillation from VFMs fed with LiDAR images can be beneficial. The motivation is three-fold: datasets with annotated LiDAR data are rather scarce since they are costly and time-consuming to obtain, the rapid development of LiDAR technology makes existing datasets obsolete quickly, and perception solutions face the challenges of representation learning for complex environments in real world applications of autonomous driving. We pose two research questions: Can the semantic knowledge learned by VFMs from natural RGB images be transferred to LiDAR representations even when no RGB/camera data are available during the LiDAR pretraining process? And if so, what knowledge is actually transferred across this large modality gap? After full fine-tuning, SSL initialization matches the downstream semantic segmentation score of an ImageNet-21k-initialized supervised baseline with the same architecture, and VFM distillation yields substantially higher linear-probe scores than MAE. Our results reveal potential in LiDAR image-based self-supervised representation learning for unlabeled long-range, high-resolution sensors, paving the way for scalable learning for modern LiDAR systems. Our code will be publicly released on GitHub after the paper is accepted for publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.