The Curation Paradox: Cleaner Labels can be harder to correct
Abstract
Cleaner labels and broader coverage are appealing goals for benchmark curation, but neither guarantees better model selection. We study speaker verification, where evaluation chooses both a model and an acceptance threshold. Agreement filtering can tie surviving label errors to model scores, violating the independence required by class-confusion correction. We express the remaining cost distortion as a covariance residual and test its consequences using frozen public speaker models with simulated label errors. An oracle intervention redistributes errors while preserving the cohort and error counts, isolating the role of error placement. Consistent identity merges and splits also produce cases where filtering improves labels but worsens the selected decision, alongside cases where filtering helps. For developers with only provisional labels, a separate budget study compares uniform, agreement, and score-stratified sampling. Broader observable coverage does not reliably improve decisions; closer agreement with the full provisional benchmark can preserve its errors. Moreover, selecting fewer pairs may save little encoder computation because pairs share recordings. These controlled results motivate evaluating curation by decision quality, information requirements, and processing cost, rather than treating label agreement or subset size as quality certificates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.