What Can Time-Series Foundation Models Actually Do? Towards Capability-Level Evaluation
Abstract
Time-series foundation models (TSFMs) achieve strong zero-shot forecasting performance, but why do they work, and what abilities support their predictions? Existing evaluations compare forecasting accuracy across datasets or under selected forecasting conditions, leaving unclear which abilities models possess and when those abilities fail. We investigate these questions through six forecasting abilities: Structure Induction, History Retrieval, State Tracking, Uncertainty Modeling, Channel Modeling, and Exogenous Modeling. For each ability, we construct controlled experiments that vary the information available to a model or the difficulty of using that information while holding other conditions fixed. We evaluate 21 frozen TSFMs through their supported interfaces, examining whether they can infer temporal structure, retrieve historical information, track changing states, represent uncertainty, combine information across channels, and respond to external conditions. The results reveal ability differences that aggregate forecasting rankings do not capture. Separating periodic structure from residual variation improves forecasting, whereas supplying additional informative channels does not consistently reduce error. Stochastic generation also does not ensure that predictive distributions express the uncertainty present in the target. These findings connect forecasting performance to how models represent, combine, and use predictive information. They motivate evaluating TSFMs through explicit ability profiles and examining whether model designs translate available information into accurate predictions and calibrated uncertainty. All code is available at https://anonymous.4open.science/r/sortinghat-32B3/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.