Contrastive Learning Across Samples and Features Through the Lens of Mutual Information
Abstract
Contrastive representation learning typically contrasts corresponding samples across two views, as in InfoNCE. However, the mutual information bound of sample-wise InfoNCE is limited by the number of negative samples and does not explicitly discourage redundancy across features. To address these issues, we study sample-wise and feature-wise contrastive learning through a common mutual information formulation. Our analysis of the feature axis reveals an additional penalty arising from the dependence and non-exchangeability of feature channels and motivates feature standardization before contrastive matching. The two objectives are complementary: the sample-wise objective distinguishes inputs, while the feature-wise objective promotes informative, non-redundant features. We prove that the mutual information captured along either axis lower-bounds the mutual information between the full embedding matrices, providing a principled basis for combining the two objectives. Our extensive experiments across image representation learning and cross-modal retrieval demonstrate that the combined objective improves over either objective alone and achieves superior performance compared with existing self-supervised learning methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.