acceptodds
Under review as a conference paper at ICLR 2027

When Do Noisy-Label Corrections Help? A Counterfactual Audit

Abstract

Human annotation is imperfect, and the standard remedy corrects labels before training, yet the corrections it proposes mix repairs with damage that must be separated. Latent-shift noisy-prediction correction (LSNPC) repairs a prediction by moving latent variables rather than the input. Their reliability can be assessed before retraining by scoring each proposed correction. First, we develop an audit with label-free scores inspired by counterfactual properties; label-free refers to scoring and selection, independently of corrector training. We apply LSNPC to training labels and study latent paths that change the decoded label as counterfactual candidates. Second, we find that confidence, one candidate score, is anti-informative where corrections rarely succeed and informative where they usually do. Score informativeness varies with the recovery rate and encoder. In paired held-out diagnostic runs, label-free proximity retains about half of its clean-label-referenced separation over chance. Across our end-to-end study, a rank threshold over the scores reverts the corrections it scores lowest before retraining. At the operating coverage the thresholded stream falls at most 0.1 accuracy points below the original noisy stream in the reported means over corrector seeds. The pattern holds across the three downstream models. At wide coverage it outperforms Co-teaching, the in-training baseline, on the crowd-labelled text streams. Its margin over the original stream is widest on the image streams. At matched coverage, label-free selection performs similarly to random rejection, identifying score discrimination as a target for improvement. This study connects counterfactual properties to selective acceptance of label corrections.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.