ShapeDINO: Scaling Self-Supervised Representation Learning for 3D Shapes
Abstract
Large-scale self-supervised learning has enabled general-purpose representations in language and vision. We investigate how far this paradigm can advance reusable representations of 3D shapes. We introduce ShapeDINO, a family of self-supervised point cloud encoders built on a standard vision Transformer backbone, ranging from Tiny to a billion-parameter Giant. ShapeDINO combines global and patch-level self-supervised learning to jointly learn object-level semantics and dense local features. To support pretraining at scale, we curate a dataset of 1.5 million 3D assets through quality filtering, geometric deduplication, and semantic rebalancing. We further distill the largest model into smaller variants to accommodate different computational budgets. Under frozen-feature evaluation, ShapeDINO achieves state-of-the-art performance on multiple classification and part-segmentation benchmarks, matching or surpassing the evaluated weakly supervised baselines. Additionally, we demonstrate that ShapeDINO can be seamlessly integrated into downstream tasks such as representation alignment, point segmentation and completion, yielding consistent performance improvements. These results demonstrate that large-scale self-supervised learning can yield reusable 3D representations that benefit diverse downstream tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.