Toward Unified Multimodal Representation Learning: Contrastive Alignment of Text, Image, and Point Cloud
Abstract
Contrastive Language-Image Pre-training (CLIP) has shown impressive performance in aligning visual and textual representations. Recent works have extended this paradigm to 3D vision, where a common strategy is to align point-cloud representations with text and image features through pairwise contrastive objectives. However, such pairwise formulations optimize modality pairs separately rather than multimodal tuples jointly. In this paper, we propose Contrastive Tensor Pre-training (CTP), a joint contrastive learning framework that extends the pairwise similarity matrix to a multimodal similarity tensor for text–image–point-cloud alignment. We further study two key aspects of joint multimodal learning. First, we show that a simple geometric scorer can effectively align the three modality representations. Second, we identify partial-match negatives introduced by joint tuple construction and show that masking them improves joint alignment. Experiments on real-world automotive LiDAR from nuScenes and dense point clouds from ULIP-ShapeNet show that CTP outperforms a matched pairwise baseline in both zero-shot classification and instance-level retrieval. The learned representations also transfer to KITTI and Waymo without target-domain training, further supporting the cross-dataset generalization of CTP.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.