A Committee of One: Detecting and Correcting Label Noise with Test-Time Augmentation
Abstract
Label noise can substantially degrade the stability and generalization of deep classifiers. Rather than designing a noise-robust training procedure, we address the problem at the data level by identifying and correcting corrupted labels before training the final model. We propose a data-cleaning framework based on a Siamese network and test-time augmentation. The network is trained once on the noisy dataset using a joint contrastive and classification objective and is then frozen. For each training sample, we generate views from a fixed set of twelve augmentation families and measure how often their predictions disagree with the recorded label. Our key insight is that mislabeled samples are more sensitive to small input perturbations, particularly near decision boundaries, whereas clean samples with larger margins remain more stable. We use this augmentation-induced disagreement to identify suspicious samples, relabel them to the most frequently predicted class when predictions are sufficiently consistent, and remove them otherwise. We further show that the gap in expected disagreement splits into a margin term, non-negative whenever clean samples dominate in sensitivity-normalized margin, and an excess we measure to be positive. The detector is trained once and never retrained; only threshold selection trains a downstream classifier. Across CIFAR-10, CIFAR-100, and Fashion-MNIST with synthetic instance-dependent noise, our approach achieves detection AUROC values ranging from to . On real annotation noise, a standard classifier trained on our cleaned labels reaches accuracy on CIFAR-10N, compared with for the strongest of nine single-model robust-training baselines evaluated together in a published common-codebase study, while matching its performance on CIFAR-100N. On Animal-10N, it achieves , compared with for the strongest of the baselines we compare against.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.