Rethinking Multimodal (Mis)Alignment for Vision-Language Representation Learning
Abstract
Vision-language models (VLMs) commonly learn by aligning image and text representations, achieving strong performance on zero-shot classification and image-text retrieval. However, this paradigm is sensitive to imperfect image-text correspondence and can overlook information that is complementary rather than shared across modalities. We investigate whether multimodal representation learning can provide a more robust alternative while retaining strong cross-modal representations. For the first time, we systematically assess CoMM as a vision-language representation learner, studying its behavior across image-text alignment quality and pretraining scale. We find that CoMM is substantially more robust to misaligned pretraining pairs than alignment-based objectives such as CLIP, while learning strong visual and multimodal representations. At larger scale, however, CoMM lags behind CLIP on standard cross-modal tasks, revealing a trade-off between robustness to misalignment and efficient learning of explicit image-text alignment. To bridge this gap, we introduce CoMM-VL, which augments CoMM's multimodal objective with an explicit image-text alignment loss within the same fused representation space. CoMM-VL substantially improves zero-shot classification and image-text retrieval over CoMM and outperforms established alignment-based VLMs, including CLIP, SigLIP, and TIPS, while preserving CoMM's robustness to misaligned data and strong multimodal representations. Across small- and large-scale vision-language datasets, these results show that explicit cross-modal alignment and multimodal interaction can be complementary rather than competing objectives. Code and pretrained weights will be open-sourced.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.