UPCDD: Unsupervised Point Cloud Dataset Distillation
Abstract
Dataset distillation (DD) aims to condense a large real dataset into a compact synthetic set that preserves its training utility while reducing storage and repeated training costs. While recent approaches have extended DD to 3D point clouds, they predominantly rely on ground-truth semantic annotations, which are difficult and expensive to acquire in the 3D domain. This dependence prevents unlabeled point-cloud data from being directly distilled and is particularly restrictive for tasks requiring costly fine-grained annotations, such as part segmentation. To address this limitation, we propose UPCDD (Unsupervised Point Cloud Dataset Distillation). To the best of our knowledge, UPCDD is the first framework to enable label-free dataset distillation for point clouds. UPCDD leverages self-supervised representation learning and adaptive feature-space clustering to discover latent semantic partitions, effectively substituting manual labels with proxy supervision. UPCDD can be instantiated by adapting existing label-dependent 3D distillation methods, while a novel augmented objective is designed to stabilize label-free distillation, facilitating effective adaptation and transfer. Experiments on 3D shape classification and part segmentation demonstrate that UPCDD retains useful semantic information under severe compression and transfers across distillation methods, architectures, and datasets. Moreover, the distilled dataset can be reused for downstream training without repeatedly accessing the full unlabeled collection, providing a compact and architecture-flexible alternative for exploiting large unlabeled point-cloud datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.