Beyond Forecast Accuracy: How Time-Series Foundation Models Respond to Longer Histories
Abstract
Longer histories can stabilize forecasts without correcting their average response to predictive information. We study this distinction in six frozen time-series foundation model checkpoints on Gaussian processes with known next-step distributions. At 2048 observations, the five deterministic checkpoints respond too weakly to the true predictive mean on average. Their excess mean squared errors are 24–45% of predictable variance, compared with 6.5% for a fitted AR(64) baseline. Yet Chronos-Bolt and Moirai vary less across histories with the same predictive mean and last observation than that baseline. Linear regression exposes the response bias. Conditional resampling then separates two sources of the regression residual: a curved average response and variation across matched histories. Exploratory bounds suggest that most residual energy comes from the latter. History length also changes the association with the last observation at a fixed predictive mean. Following exploration, a prospective three-lag test supports opposite Bolt–Moirai changes from 512 to 2048, averaged over 16 conditions; Moirai's last-observation response varies nonmonotonically with history length. These comparisons show why greater forecast stability need not repair the main sources of error.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.