acceptodds
Under review as a conference paper at ICLR 2027

Trust No Single Noise Model: Certified Active Dataset Repair

Abstract

A recorded label is an imperfect observation of the label we want, separated from it by a corruption process that recurs across records and can be modelled, so that a small audit budget repairs some records directly and, through the model, others indirectly. But the noisy labels alone do not identify that process: several models fit the same data and prescribe different audits and repairs, and existing methods commit to one and offer no guarantee when it is wrong. We treat the problem as a game: nature picks one corruption model from a declared set of candidates; the learner chooses audits sequentially and then corrects the unaudited labels. Solving the game yields a repaired dataset with a certificate, the largest error expected to remain under any candidate, computable without gold labels. The minimax value is the Bayes value under a least-favourable weighting of the candidates, which makes the game solvable by dynamic programming, and choosing audits one at a time can be arbitrarily worse than planning the sequence. On the CIFAR-10H test set the minimax policy attains the lowest certified risk at every budget among 23 arms, including eight published cleaners: every cleaner's certified risk stays above 5% while ours falls from 5.4% to 0.9%, at a price of about one point of true error against a Bayesian weighting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.