Multi-Source Multi-View (MSMV): A robust SSL framework for imbalanced data
Abstract
One of the most challenging problems in self-supervised learning (SSL) is to achieve robustness of the learned representations irrespective of whether the data is balanced or imbalanced. SSL methods address this problem through techniques like feature reweighting, regularization or by making use of out of distribution(OOD) data. However, contrastive self supervised learning (CSSL) is more sensitive to imbalance, as frequent samples produce disproportionately more contrastive pairs leading to biased representations. Our analysis on the long-tail behavior in SSL clearly shows that most existing methods are biased in terms of *gradient ratio* which could lead to degradation of representation quality as well as poor decision boundary separation. We believe that the core reason for such quality degradation comes from the traditional *multi-view* design on which most existing SSL frameworks are based. Building on these findings, we design a novel contrastive learning framework (MSMV) in which multiple sources are considered to promote balanced information flow across views rather than a single source as in multi-view learning. The proposed model is enhanced by introducing a new contrastive loss function that helps in better representation learning for tail classes. Additionally, we adopted a unique augmentation strategy to improve the overall performance. Extensive experiments achieve consistent improvements across multiple benchmarks, including a 5% increase in linear evaluation and an 8% increase in few-shot accuracy on Cifar100-LT with ResNet-18, highlighting its effectiveness in imbalanced scenarios. We also find consistent gains on other datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.