Auditing Common-Threshold Extrapolation for Rare Events: Identifiability, Prediction, and Failure Modes
Abstract
When can common-threshold labels improve estimation of a rare-event probability? We develop a target-specific audit that distinguishes identifiability, statistical efficiency, and learning performance. For a locally identifiable target, the endpoint-to-common efficient variance ratio is , where is the target log-probability gradient and is common-label Fisher information. A target direction outside the Fisher range rules out regular root- estimation. Spectral diagnostics and a perturbation bound assess the stability of estimated information. In a two-rate benchmark, common-only extrapolation improves deep-tail accuracy, misspecification reverses the gain, and an added threshold restores identifiability. A prospective neural study evaluates five targets at three model sizes with approximately matched parameter counts. A pointwise Fisher score predicts the mean direction of the error difference in all 12 cells outside its indifference band, with rank correlation 0.917 across all 15 cells. At the deepest target, endpoint-to-common error ratios are 1.67–1.74. An independent 20,000-observation common-label pilot identifies harm at three targets and abstains at two, falling short of its pre-specified decision-coverage criterion. Temporal tests show mixed effects on USGS earthquakes and 3.5–16.8% lower held-out log loss in four gas-turbine target–year cells. The resulting approach connects label design to testable learning predictions through target identifiability and information margins.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.