acceptodds
Under review as a conference paper at ICLR 2027

Not All Mislabels Are Equal: Efficient Task-Aware Auditing with Confidence Sequences

Abstract

We study efficient, adaptive auditing of categorical labels in a fixed dataset, motivated by biodiversity-credit programs in which expert verification of species records is costly and downstream scores (e.g., payment) computed from corrected labels are more sensitive to errors on certain labels than on others. Our framework couples anytime-valid inference with a task-aware adaptive auditing policy, representing label corrections through a confusion matrix and prioritizing which data points to inspect next based on the uncertainty within their class labels and the corresponding influence of these labels on the downstream score function. For linear scores such as the mislabel rate, our priority-aware approach certifies the overall label error rate to +/- 2 percentage points with 24% and 11% fewer audits compared to random sampling on biodiversity data from Kenya and a widely-used CIFAR-10N benchmark, respectively. For nonlinear scores such as the Hill–Shannon diversity, we exploit a linear-envelope representation that reduces inferences to a family of linear-score audits; priority-aware sampling again reaches the target precision with 36% fewer audits. When the cost of distribution-free guarantees becomes prohibitive, we develop a model-based Bayesian alternative that trades assumption-free validity for considerably fewer audits. Together, these results show task-aware auditing can significantly reduce the number of expert audits needed while providing finite-sample, anytime-valid guarantees.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.