acceptodds
Under review as a conference paper at ICLR 2027

On the Origin of the Modality Gap: A Mechanistic Perspective via Correspondence Noise

Abstract

Contrastive vision-language models exhibit a persistent modality gap despite being trained to align images and text, yet its underlying mechanism remains unclear. We provide a mechanistic account demonstrating that a small number of high-offset dimensions disproportionately shift the ground truth retrieval rank of samples with weak, ambiguous, or erroneous image–text correspondence through their cross-modal interaction. Controlled perturbations of correspondence quality reproduce this behavior, providing causal evidence that correspondence noise contributes to the emergence and function of the gap. We introduce a correspondence noise score, show consistency across checkpoints and diverse pretrained models, and demonstrate that filtering correspondence noise before training jointly reduces the modality gap and improves downstream performance. Finally, we show that gap magnitude is positively but inconsistently associated with performance across converged models, making it an unreliable proxy for representation quality, while its functional importance is primarily geometric rather than semantic.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.