Self-Supervised Point Cloud Dataset Distillation for Efficient Representation Learning
Abstract
Dataset distillation (DD) aims to replace a large training dataset with a compact synthetic set that preserves its training utility, substantially reducing repeated storage and training costs. However, existing point-cloud DD methods are developed under supervised settings and rely on semantic annotations to construct distilled data. While self-supervised learning (SSL) makes large unlabeled point-cloud collections usable without annotation, repeatedly training on the full dataset remains computationally expensive. Moreover, a pretrained checkpoint merely provides model-specific rather than data-level reusability. This motivates a fundamental question: can the representation-learning utility of a massive unlabeled point-cloud dataset be distilled into a small, reusable synthetic set? We study self-supervised dataset distillation for point clouds and propose a distillation framework for constructing compact training sets from unlabeled 3D data. Rather than relying on semantic classes or discrete pseudo-labels, we formulate distillation around self-supervised representation learning and optimize synthetic point clouds according to their ability to train newly initialized encoders. Specifically, our proposed framework leverages masked predictive learning to preserve informative global and local structures under a highly constrained synthetic budget. We evaluate the distilled synthetic sets on multiple standard point-cloud benchmarks via frozen representation evaluation against equally sized real-subset controls. Extensive experiments demonstrate that the optimized synthetic data yields non-trivial representation-learning utility compared to equally sized real subsets across multiple compression regimes. These results demonstrate the potential of self-supervised point-cloud dataset distillation as a data-level mechanism for compressing and reusing large unlabeled 3D collections.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.