acceptodds
Under review as a conference paper at ICLR 2027

The Price of Choosing: Forecast Selection Under Fresh Calibration

Abstract

When a forecast chosen with one calibration pool is calibrated again before deployment, the two calibrations need not rank candidates alike. We make the expected interval score after a fresh, finite calibration pool the comparison target. An exact construction shows that reusing one pool can select the worse forecast regardless of the number of queries. For A fixed predictors and N observations, rank weights and two hinges compute the complete-subset risk estimate and every fixed-pool-size delete-one value exactly in O(AN log N) arithmetic operations; across 12 settings, this is 1.44–24.69 times faster than a custom batched dense reference computing the same outputs. For globally nearby endpoints with a common bounded scale, calibration ranks give a paired score envelope without bounding the response, and its uniform constant 8 is sharp across ranks. Label-cost bounds then separate exact assessment from reliable selection. In a 72-setting unbounded-support study, the envelope enables beneficial controlled updates in up to 20 settings, while a query-level bound never updates. On nine public forecasting datasets, calibration removes most of a matched fitting gain, from 5.34 to 1.52 percent, and from 7.05 to 3.13 percent over ten seeds; in a prespecified external test, the seven hourly datasets, one of its two primary groups, retain 62 percent of the gain. Pre-calibration comparisons can therefore overstate the surviving difference. Candidate forecasts should be compared by their risk after the calibration they will actually receive, and its exact estimate is now cheap to compute.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.