Forecasting the Forecasters: Training-Free Two-Phase Aggregation of Frozen Time-Series Foundation Models
Abstract
Frozen time-series foundation models differ in accuracy across datasets and over time, making it difficult to decide which forecasts to trust. Combination weights fixed from historical backtests cannot adapt to new outcomes, while deployment-only aggregation has little feedback when few forecasts are scored. Learned routers require additional training. To address these limitations, we introduce TIMEHEDGE, a training-free aggregation procedure that combines unscored historical backtests, called practice, with deployment feedback. Practice sets initial weights from pooled and series-specific errors; deployment feedback updates them. The weights combine forecasts at each quantile level, followed by sorting. We extend classical expert-path analysis to this two-phase protocol and ordered quantile levels. An exact identity quantifies how practice changes a deployment regret bound, where regret measures cumulative excess loss relative to a sequence of model choices. Convexity and rearrangement bound forecast pinball loss by weighted member losses. On GIFT-Eval, TIME and fev-bench, TIMEHEDGE outperformed all 22 tested same-pool baselines: eight individual models and 14 selection or aggregation methods. These gains held for both mean absolute scaled error and continuous ranked probability score on all three benchmarks. Ablation studies indicate that practice drives most of the gains: pooled and series-specific evidence each help, and the same protocol benefits other online learners.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.