Can Time Series Foundation Models Scale at Test Time?
Abstract
Test-time scaling aims to improve forecasts through additional inference computation, yet its benefits beyond strong raw-input baselines remain unclear. We study four frozen time series foundation models (TSFMs) on univariate point-forecasting tasks from eight datasets, distinguishing fixed-input sampling, multi-view inference, and history-based selection and aggregation. Repeated sampling improves forecasts with diminishing returns: on eight designated series, doubling the sample count from 64 to 128 reduces average relative MAE by less than 0.6% for each stochastic model. Multi-view aggregation can provide further gains. Across 48 additional series from three datasets, Moirai's historically selected multi-view ensemble achieves a dataset-balanced mean MAE reduction of 7.42% against a historically selected raw-input ensemble. However, gains vary across models and tasks, and stronger raw-input controls reduce some apparent improvements. Historical improvements can reverse at test time, while candidate-pool expansion can worsen selection and incur additional calibration costs. Our findings distinguish gains from sampling, views, and aggregation, highlighting the gap between candidate potential and forecasting improvements realizable using observed history. Code is available at [https://anonymous.4open.science/r/forecasting-tts-fb18b6a3/](https://anonymous.4open.science/r/forecasting-tts-fb18b6a3/).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.