acceptodds
Under review as a conference paper at ICLR 2027

Label Provenance Learning: Modeling Where Noisy Labels Come From

Abstract

Noisy-label learning commonly represents supervision as input–label pairs, discarding the source of each label. Yet large datasets pool labels from different websites, collection protocols, or human annotators, whose errors need not follow the same pattern. We study this setting as label provenance learning (LPL): the learner observes , where is a potentially corrupted label, identifies its source, the clean label is latent, and no source is assumed trusted. Provenance can act through two channels, a routing channel describing how examples reach sources and a labeling channel describing how sources produce labels; each may or may not depend on the provenance, defining four nested prediction structures. We derive provenance-conditioned variational learning (PCVL), which fits any of these structures under a single objective, and a permutation diagnostic that selects among them before training, using no clean labels. On CIFAR-10N the diagnostic detects neither dependency, and the provenance-blind structure has the highest mean accuracy among the four; on dopanim it detects a source-dependent labeling channel, and activating only that channel improves balanced accuracy over both provenance-blind PCVL and DivideMix. The diagnostic's decisions agree with post-hoc oracle audits in every real-data setting. A controlled CIFAR-10 cross that varies the two mechanisms separately shows the labeling channel tracking its generating mechanism, with weaker evidence for routing; a matched stress test shows that shuffling provenance removes the gain. Provenance is therefore worth modeling when its conditional role is supported, rather than activated by default.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.