acceptodds
Under review as a conference paper at ICLR 2027

Point-S2AE: Structural and Semantic Masked AutoEncoders for Point Cloud Pre-training

Abstract

Masked autoencoders have been widely explored in point cloud self-supervised learning, whereby the point cloud is generally divided into visible and masked patches for reconstruction. These methods typically reconstruct masked patches independently based on visible representations, driving the encoder to learn local geometric structures. However, they treat the reconstruction of each masked patch as an isolated task, neglecting the inherent structural dependencies between adjacent patches. Furthermore, their strict reliance on local geometric reconstruction often fails to capture the holistic semantic consistency of 3D objects, limiting the robustness of the learned representations. To overcome these limitations, we propose a novel pre-training framework that augments the standard Masked Point Modeling (MPM) paradigm with two simple yet effective components: Latent Autoregressive Alignment (LAA) and Global Self-Distillation (GSD). Specifically, LAA explicitly models the spatial transitions between masked patches by enforcing an autoregressive predictive relationship between shallow and deep latent representations in the decoder, utilizing an InfoNCE contrastive loss to prevent representation collapse. Concurrently, GSD extracts consistent global semantics by applying a bidirectional self-distillation loss on the class tokens from two distinct masked views of the same point cloud in the encoder. Extensive experiments on standard benchmarks demonstrate that Point-S2AE significantly enhances the learned representations while maintaining high pre-training efficiency. Notably, Point-S2AE achieves great improvement over the Point-MAE baseline, particularly surpassing it by 5.16% on OBJ_BG, 5.34% on OBJ_ONLY, and 5.35% on PB_T50_RS for 3D object classification on the ScanObjectNN dataset, and establishing new state-of-the-art results on ModelNet40 few-shot learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.