Detection Is Not Repair: Evaluating Budgeted Label Verification for Group Robustness
Abstract
Does uncovering more label errors necessarily improve worst-group accuracy after retraining? In budgeted data cleaning, standard practice prioritizes detection precision—identifying as many corruptions as possible within a fixed query budget. In this work, we investigate this premise when group annotations are inaccessible during data acquisition, and uncover a fundamental decoupling between error detection and downstream group robustness. Across extensive empirical evaluations on vision and text benchmarks, we document widespread empirical inversions where acquisition policies that correct substantially fewer errors achieve markedly superior worst-group accuracy compared to high-precision detectors. This divergence arises because high-precision selection systematically starves underrepresented, vulnerable subgroups of corrections, skewing the retrained decision boundary unfavorably. Mathematically, we formalize this structural failure: without group annotations, even an idealized oracle detector incurs a worst-case regret approaching 50 percentage points due to the inherent non-identifiability of latent group structure. Finally, we explore practical remedies in the absence of group metadata, showing that simple observed-class quotas safeguard against decision boundary destabilization, while tracking post-retraining clean validation improvement guides policy selection in balanced vision domains. Our results reposition data cleaning from passive noise filtering to an active intervention on classifier decision boundaries, underscoring that label verification must be evaluated by downstream subpopulation robustness rather than raw detection metrics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.