Hardness Is Not Enough: What Contrastive Error Audits Identify
Abstract
Contrastive error auditors rank the items of a data pool by their resemblance to synthetic corruptions, and harder, more plausible corruptions are generally preferred. We prove that no criterion computed from population testing difficulty can certify how well such an auditor ranks real errors. With the pool, error prevalence, representation and real error process held fixed, two corruptions can share the complete binary testing region, and hence the optimal ROC curve, every binary Bayes risk and every -divergence, while their optimal scores rank the real errors optimally and in exact reverse. This holds on every atomless pool with a nonconstant error posterior, at every continuous difficulty level, and an exact criterion decides it on finite pools. Fidelity, in contrast, is sufficient: audit regret is bounded by the total-variation distance between the corruption and the error distribution plus the contrastive AUC shortfall. Crossed detector interventions on knowledge graphs, including a prespecified replication in an independent architecture, show that the preferred training corruption depends on the errors evaluated, and changing only those errors reverses the order of two published detectors on FB15k-237, by 0.008 in top-5% precision. The dependence persists in image audits at matched measured difficulty and in human-annotated summaries. Synthetic corruptions should therefore be justified against the target error process and evaluated on held-out target errors alongside controls fitted on the same labels.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.