Beyond Sample Reliability: Relation Identity Determines Noise Harm in Supervised Contrastive Learning
Abstract
In supervised contrastive learning (SupCon), noisy labels alter not only individual supervision but also cross-sample positive relations. We ask which of these relations should be trusted. Through controlled noise placement and quantity-matched interventions, we find that sample difficulty can help identify corrupted supervision but does not consistently predict its representation harm, while changing which relations are modified can produce substantially different outcomes under the same intervention budget. Motivated by this observation, we propose Posterior-Guided Relation Correction (PRC), which learns class posteriors from noisy data and uses their pairwise agreement to reallocate normalized positive target mass among observed-positive relations while preserving instance positives. We further characterize this reweighting at the gradient level: the effect of a relation perturbation depends jointly on its assigned target mass and the current gradient geometry, and posterior-guided reweighting is locally beneficial only when the support score aligns with relation utility. Experiments across synthetic corruption, varying noise rates, and real-world noisy-label benchmarks show consistent representation improvements. Further analysis reveals a clear boundary: strong relation diagnosis does not necessarily support reliable intervention, especially when recovering missing positive relations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.