When Better Models Score Lower: Test-Set Label Noise Inverts Method Rankings in Learning with Noisy Labels
Abstract
Label noise in fine-grained recognition is not random: annotators confuse semantically adjacent classes, so errors concentrate on twin class pairs. This affects both the methods of learning with noisy labels (LNL) and the protocol used to compare them. (i) Evaluation. Test sets are annotated by the same procedure as training sets, so they carry the same structured noise, and a noisy test set does not merely lower scores: it inverts rankings. Measured accuracy obeys an exact two-term law, , where NF ("noise fitting") is a method-intrinsic propensity to agree with the noise direction and is anti-correlated with true quality. Across three fine-grained benchmarks and seven methods, 58 of 63 method pairs change order, with a median crossover at ; a symmetric-noise control shows NF at chance level and no inversion. The inversion is not an artefact of the injected noise: it also appears on real crowd re-annotation of CIFAR-100N, where no CLIP output enters the noise at any point and a method the clean test set rejects by a wide margin comes out ahead on the human-annotated test set of the same images. The closed-form crossover recovers, as its small- limit, the earlier observation that test-set errors can flip model selection. Nor is the law tied to our method, backbone or modality: it holds on a 1000-class human re-annotation of ImageNet val, where with no clean labels the correction still recovers the human-verified accuracy to within a point, and the rankings invert again on 20 Newsgroups with a bag-of-words model and no pretrained network. Given any class-level confusion partner map it predicts the accuracy curve of any model, with NF the only free quantity; and the threshold grows as when only a fraction of test errors shares the training direction. (ii) Methods. To measure this we need a method that is actually strong under semantic-confusion noise, so we build SASR — a vehicle for the measurement, not a claimed advance. Every remedy based on data-internal statistics fails here: kNN label propagation inverts and lands below the CLIP zero-shot baseline. SASR beats cross-entropy by 14–18 points on all three benchmarks. Against contemporary CLIP-based remedies the result is less flattering, and we report it: a plain CLIP hard-selection baseline matches SASR on CUB-200 and beats it on Cars196. (iii) Protocol. Correcting an accuracy estimate for an imperfect evaluator is classical, and the standard correction assumes that model errors and annotation errors are independent. LNL methods violate that assumption by construction, being optimised on the same noise that reappears in the test set. PAD (Pair-Agreement Debiasing) replaces the assumption with a measurement: using class-name semantics to obtain the noise direction, a single model's clean accuracy is recovered in closed form from two agreement rates on one noisy test set (error for ). PAD reduces exactly to the classical correction when independence holds, and as the two unknowns become unidentifiable — the distortion is worst exactly where it cannot be repaired. PAD needs no clean reference, no retraining and no model internals, so it can be applied post hoc to an existing leaderboard.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.