Prevalence Determines Precision: Silent Contamination in Detector-Defined Datasets
Abstract
A large share of machine learning datasets are not observed but constructed: a detector, heuristic, or model is run over a pool of candidates, and whatever it accepts becomes the dataset. Weak and distant supervision, pseudo-labelling, event extraction, and most anomaly-detection benchmarks all have this form. The precision of the resulting dataset is not a property of the detector. It is governed by the prevalence π of true positives in the pool the detector is deployed on, through elementary Bayes. This is textbook, and it is nonetheless almost never measured end-to-end, for a structural reason: from inside a detector-defined dataset the false positives are unidentifiable, so the quantity that matters cannot be estimated by the people who need it. We report a case in which it can be measured. For the same instrument and period we hold both a detector-defined event dataset and an independent official index that reveals, for every detected item, whether it is real. One detector, three candidate pools, three datasets: phantom rates of 81.7%, 9.0%, and 0.0%. A practitioner transferring precision from the two high-π pools to the low-π pool predicts 0.955 against a measured 0.183, an error of +422%; the Bayes expression predicts within 3.3% across all three pools. Beyond confirming the mechanism we report three findings that we believe are new. First, the detected dataset’s response curve is an exact convex combination of a true-event and a phantom component (identity residual 1.1 × 10−16), with phantoms outnumbering true events 473 to 308 — contamination is not noise added to signal but a second signal with its own shape, inherited from the detector’s acceptance rule. Second, the direction of contamination is a property of the estimator, not of the data: on identical windows, one statistic shows contamination diluting the effect and another shows it inflating the effect, because the second statistic’s denominator is itself contaminated. Third, a normalisation in common use turns the estimator into a mean of ratios whose expectation does not exist; on the same 335 events it returns 0.40 where the well-defined estimator returns 0.10.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.