On the Origin of the Modality Gap: A Mechanistic Perspective via Correspondence Noise
Abstract
Contrastive vision-language models exhibit a persistent modality gap despite being trained to align images and text, yet its underlying mechanism remains unclear. We provide a mechanistic account demonstrating that a small number of high-offset dimensions disproportionately shift the ground truth retrieval rank of samples with weak, ambiguous, or erroneous image–text correspondence through their cross-modal interaction. Controlled perturbations of correspondence quality reproduce this behavior, providing causal evidence that correspondence noise contributes to the emergence and function of the gap. We introduce a correspondence noise score, show consistency across checkpoints and diverse pretrained models, and demonstrate that filtering correspondence noise before training jointly reduces the modality gap and improves downstream performance. Finally, we show that gap magnitude is positively but inconsistently associated with performance across converged models, making it an unreliable proxy for representation quality, while its functional importance is primarily geometric rather than semantic.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.