GradMix: Improving Supervised–Self-Supervised Compatibility through Attribution-Guided Feature Coverage
Abstract
Self-supervised learning (SSL) has been widely applied to supervised models, either during pre-training or as an auxiliary objective, to improve representation feature coverage beyond the information directly associated with class labels. This is particularly relevant to open-set recognition (OSR), which requires models to discriminate among known classes while retaining information beyond class-discriminative labels to distinguish unknown samples. However, when applied as an auxiliary objective for OSR, SSL can lower known-class classification accuracy, leading to a generalization and discrimination dilemma. To understand this problem, we analyze the representation-learning characteristics of supervised and self-supervised learning and attribute the degradation to conflicting representation-learning signals. We then propose GradMix, a data augmentation strategy that enables the SSL objective to learn more generalizable representations, letting SSL learn a broader range of representations compatible with supervised learning, thereby expanding the solution space shared by both objectives. Experimental results on standard and domain-specific OSR benchmarks show that GradMix consistently improves over naive joint supervised–self-supervised training and achieves competitive performance against state-of-the-art OSR methods. Beyond OSR, evaluations on out-of-distribution detection, robustness to common corruptions, and linear probing of SSL representations further demonstrate its ability to improve both representation generalization and discriminative precision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.