PATHTAIL: Calibration Attribution from Finite Forecast Samples
Abstract
Calibration audits often estimate a specified prediction error from a finite set of forecast trajectories. When that error is defined by a population projection, fitting the projection on the forecast sample can change the audit target. We give an exact counterexample in which the intended residual component is zero, yet a residual fitted to only eight forecast paths has nonzero mean. PATHTAIL constructs a score and companion statistic whose expectations preserve the original residual component and its tolerance. On finite domains, we characterize the minimum fixed forecast budget required by bounded exact scores and show that it depends on the target even within a fixed nuisance space. These scores support anytime-valid sequential auditing. A paired study further shows that exact score choice matters: a pooled-count construction substantially improves detection under rare-state predictive laws while losing under uniform laws, demonstrating that maximizing the mean signal does not maximize detection. Controlled feedback from four native forecasting models across eight data sources verifies target preservation, while a replayable count archive exposes the remaining limits of detection and sampling efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.