Ada3D: Adaptive Cross-Dimensional Alignment for Universal 3D Representation
Abstract
Motivated by the significant performance of contrastive learning models in vision-language tasks, recent efforts have extended such architecture into 3D scenes for richer representation. However, existing approaches typically utilize the pretrained 2D ViT models to encode original 3D group features, which overlooks the inherent domain gap between these two modalities and leads to suboptimal performance. Observing that 3D point clouds and 2D images share geometric correspondences that can be explicitly correlated through cross-dimension projection, we propose Ada3D, a framework that focuses on the adaptive transfer from the 3D to 2D feature spaces to enable universal 3D representation. We achieve this transfer capability via a dual mechanism comprising a pre-adapter and a post-adapter. The pre-adapter bridges unordered point cloud groups and regular image patch sequences through the point-level and patch-level fusion paradigm. Subsequently, the post-adapter utilizes a mix-attention mechanism to capture the latent features in the refined 3D groups and transforms them into the ViT input space. Furthermore, we introduce a specialized Gaussian masked autoencoder and a shallow Transformer layer finetuning strategy to preserve the generalization and reduce computation. Extensive experiments demonstrate that our approach achieves remarkable results on diverse 3D universal representation tasks, including zero-shot open-vocabulary classification, segmentation, and visual question answering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.