acceptodds
Under review as a conference paper at ICLR 2027

Ada3D: Adaptive Cross-Dimensional Alignment for Universal 3D Representation

Abstract

Motivated by the significant performance of contrastive learning models in vision-language tasks, recent efforts have extended such architecture into 3D scenes for richer representation. However, existing approaches typically utilize the pretrained 2D ViT models to encode original 3D group features, which overlooks the inherent domain gap between these two modalities and leads to suboptimal performance. Observing that 3D point clouds and 2D images share geometric correspondences that can be explicitly correlated through cross-dimension projection, we propose Ada3D, a framework that focuses on the adaptive transfer from the 3D to 2D feature spaces to enable universal 3D representation. We achieve this transfer capability via a dual mechanism comprising a pre-adapter and a post-adapter. The pre-adapter bridges unordered point cloud groups and regular image patch sequences through the point-level and patch-level fusion paradigm. Subsequently, the post-adapter utilizes a mix-attention mechanism to capture the latent features in the refined 3D groups and transforms them into the ViT input space. Furthermore, we introduce a specialized Gaussian masked autoencoder and a shallow Transformer layer finetuning strategy to preserve the generalization and reduce computation. Extensive experiments demonstrate that our approach achieves remarkable results on diverse 3D universal representation tasks, including zero-shot open-vocabulary classification, segmentation, and visual question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.