The Price of Generality in Time Series Foundation Models: GRAIN for Yearly Forecasting
Abstract
Foundation models are typically trained as generalists, optimised for average predictive performance across broad and heterogeneous tasks. Motivated by the classical James-Stein phenomenon, we develop a theory that explains why such models can remain suboptimal on specific domains despite strong aggregate performance. We interpret pretraining as empirical Bayes-risk minimisation under a corpus-induced task prior, and define the *price of generality* as the excess risk incurred on a target domain when using a rule optimised for a broader population. While the argument applies broadly to foundation models, it has practical relevance mostly for cases with a limited amount of information in the prompt. Consequently, we study yearly forecasting as a concrete application of this theory. Yearly series are practically important but systematically underrepresented in existing pretraining corpora, and their short histories make predictions more sensitive to prior misspecification. Following our developed theory, we introduce GRAIN, a specialist prior-data fitted network trained on simulations from a posterior-informed prior estimated from real yearly series. On the yearly subset of the GIFT-Eval, GRAIN outperforms generalist foundation models and statistical baselines without data leakage, supporting the benefit of aligning the training prior with the target task population.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.