Theoretical Analysis of Multimodal Contrastive Learning: When and How Hard Negatives Help
Abstract
Multimodal contrastive learning has emerged as a powerful framework for learning visual and semantic representations, yet it often struggles to capture rare features. Hard negative sampling has been empirically shown to improve fine-grained recognition, but a rigorous understanding of when and how it helps remains limited. In this work, we develop a theoretical framework to analyze the training dynamics of multimodal contrastive learning with Transformer-based encoders. Our analysis reveals an implicit bias in vanilla contrastive learning toward common features, while rare features can remain weak or even be overlooked. Motivated by this finding, we show that hard negative samples can alleviate this imbalance by suppressing common features shared across negative pairs, thereby strengthening the learning of rare features. We further characterize when hard negative sampling becomes harmful, showing that excessive use of hard negatives can overly suppress common features and disrupt the balance between common and rare feature learning. Finally, we show that a simple selection strategy based on thresholding can mitigate this effect and enable more effective use of hard negative samples. Our main theoretical findings are verified by experiments on both synthetic data and standard benchmark datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.