Certifying Data Rankings Under Deployment Uncertainty: When Benchmark Agreement Transfers and When to Abstain
Abstract
Data products are often ranked before the downstream learner is fixed. In a prospectively frozen study on untouched human-activity data, histogram boosting and a standard multilayer perceptron (MLP) significantly prefer opposite acquisition products, drawn from different subjects, under familywise simultaneous intervals. The comparison averages over the registered acquisition and training randomness while the archive, training pools, and test cohort stay fixed, so it does not claim that the reversal will recur for new subjects or sites. The result shows why a single "best dataset" is not always the right thing to acquire. We then ask when agreement among benchmark learners can still certify a ranking under deployment uncertainty. In an exact Bayesian linear-Gaussian model with a shared target, each dataset induces an information operator and each executable learner induces a trace response. Preferring one dataset for every learner in a declared family is then exactly a dual-cone order. No finite benchmark catalog covers all responses in dimension two or more. Projection onto the benchmark cone gives the exact largest normalized reversal that every benchmark accepts, and finite datasets realize every such separating direction. On the positive side, if the convex hull of the benchmarks covers the deployment family to within , a benchmark margin transfers whenever , and the auditor abstains otherwise. Learners whose response depends on the data add an exact one-sided drift penalty . In controlled families the certificate is sound and non-vacuous: a registered sign-aware rule certifies 236 of 480 ridge pairs, which is 76.1% of the 310 pairs where a conservative oracle confirms a uniform preference, with no false certificate. The natural certificate arms return no positive certificates, and a second prospective dataset is ruled uncovered because its samples miss classes. The theorem thus certifies rankings only inside an explicitly covered response family. Learner-free scores remain useful proxies, but a ranking that holds for every deployment requires a declared family, measured coverage, and honest abstention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.