acceptodds
Under review as a conference paper at ICLR 2027

Toward Unified Multimodal Representation Learning: Contrastive Alignment of Text, Image, and Point Cloud

Abstract

Contrastive Language-Image Pre-training (CLIP) has shown impressive performance in aligning visual and textual representations. Recent works have extended this paradigm to 3D vision, where a common strategy is to align point-cloud representations with text and image features through pairwise contrastive objectives. However, such pairwise formulations optimize modality pairs separately rather than multimodal tuples jointly. In this paper, we propose Contrastive Tensor Pre-training (CTP), a joint contrastive learning framework that extends the pairwise similarity matrix to a multimodal similarity tensor for text–image–point-cloud alignment. We further study two key aspects of joint multimodal learning. First, we show that a simple geometric scorer can effectively align the three modality representations. Second, we identify partial-match negatives introduced by joint tuple construction and show that masking them improves joint alignment. Experiments on real-world automotive LiDAR from nuScenes and dense point clouds from ULIP-ShapeNet show that CTP outperforms a matched pairwise baseline in both zero-shot classification and instance-level retrieval. The learned representations also transfer to KITTI and Waymo without target-domain training, further supporting the cross-dataset generalization of CTP.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.