Are Universal Time Series Foundation Models Based on a Category Error?
Abstract
Time series foundation models (TSFMs) achieve strong forecasts across diverse domains with little or no task-specific training. Yet what makes this transfer work, and why can adding data or longer histories sometimes fail to help? We investigate a possible *category error*: treating a shared numerical format as sufficient grounds for shared predictive structure and information. Time series are a container for observations from different processes; the same history can imply different futures when their generating rules remain unresolved. Under squared loss, averaging those rules suppresses conflicting components while retaining common structure, whereas informative context can make the rules distinguishable. We develop this concern into **four questions** about task symmetries, rule information, source mixing, and lightweight reuse of forecasting ability. Controlled interventions verify predicted costs of invalid sharing and show when additional information restores suppressed components. Small-scale pretraining establishes that the value of auxiliary data depends on its source and allocation. An exploratory distillation study finds that a tiny, history-conditioned operator model acquires useful predictive ability from few TSFM examples, with better performance when its training includes the target source. These findings give concrete grounds for reconsidering the assumptions behind universal forecasting and motivate models, losses, training procedures, and data construction that account for what can be shared and what must be distinguished.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.