acceptodds
Under review as a conference paper at ICLR 2027

Alignment Governs the Modality Gap: Understanding Its Interplay with Uniformity and Downstream Performance

Abstract

Contrastively trained vision-language models exhibit a persistent modality gap, with image and text representations occupying distinct regions of a shared embedding space, yet what governs this gap remains unclear. While one line of prior work has attributed the modality gap to modality-wise localization, we instead investigate the role of cross-modal alignment. Through an analysis of InfoNCE gradients, we show that InfoNCE does not explicitly promote alignment between the two modality regions. Using a controlled intervention that explicitly strengthens this form of alignment, we substantially reduce the modality gap, as measured by both positive-pair similarity and centroid distance, while largely preserving zero-shot classification and cross-modal retrieval performance. Notably, within-modality uniformity deteriorates as the gap decreases, demonstrating that improved uniformity is not necessary for gap reduction. We further uncover an empirical trilemma among positive-pair similarity, within-modality uniformity, and downstream performance: across controlled and alternative interventions, we find no examined configuration that simultaneously improves positive-pair similarity and uniformity while preserving downstream performance. Finally, we find that improved uniformity provides an additional route to reducing centroid distance without improving positive-pair similarity, yet targeting higher uniformity during training does not consistently yield a more favorable modality-gap–performance relationship under distribution shift. These findings identify direct cross-modal alignment as a mechanism governing modality-gap reduction, while revealing the metric-dependent role and limitations of within-modality uniformity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.