A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation
Abstract
Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the resulting contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration: sharp observations promote many blocks above threshold, while attenuated predictions from squared-error-type regression leave the same blocks below it. We repurpose classical monotone calibration as a symmetric audit: a post-hoc transform is fitted on held-out data and applied separately to each system. The transform requires no retraining and cannot reverse pixel ordering, so any contrast it reproduces cannot establish improved spatial ranking. On SEVIR, two released checkpoints of one architecture differ by -29.5% in extreme-threshold CSI before the control and by +5.3% after it. Across 450 pairwise contrasts among 6 systems, the difference in pooled frequency-bias deviation is associated with how far the CSI contrast moves under the control ( [+0.753, +0.832]), and 51 contrasts reverse sign. At CasCast's published extreme-event operating point, the cascade-over-backbone CSI gap falls from 0.1601 to 0.0339, a 78.8% reduction; the remaining gap stays positive. The effect persists when the transform is fitted on a window preceding the test period, and it acts in both directions: calibration also reveals advantages hidden by a better-calibrated baseline. A rarity sweep on geostationary infrared imagery shows the relative gain from the transform growing as events become rarer, crowd-density experiments reproduce the bias-gain relationship under patch-sum pooling, and semantic segmentation marks the boundary: where frequency bias is already near one, the average change is small. The confound therefore requires both a fixed operating point and a training regime that leaves the output miscalibrated there. We recommend reporting pooled frequency bias and a symmetric held-out FreeKnob Audit, released as an installable package, alongside rare-event pool-and-threshold scores.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.