Do TSFMs Suffer from Unreliable Contexts? Rethinking Context Usage in Forecasting with Time-Series Foundation Models
Abstract
Time-series foundation models (TSFMs) have recently improved frozen-backbone forecasting by scaling model capacity, pretraining corpora, and domain coverage. Unlike large-language-model inference, where prompt construction has become an explicit object of optimization, TSFM inference typically passes historical context directly to the forecaster and implicitly treats its observations as equally useful and trustworthy. This assumption can fail when the context contains stale regimes, local anomalies, misleading trends, noise, or irrelevant history. Scaling a forecasting backbone alone does not decide which evidence should be trusted for a particular instance. We therefore study evidence-aware forecasting: how should context become a selectable decision variable before a frozen forecaster models the future? The difficulty is that context reliability depends on both the instance and the forecaster, while the future outcome needed to assess an intervention is unavailable at inference. We introduce CARE-TSFM, a frozen-backbone inference-adaptation framework whose external controller is trained offline with counterfactual utility supervision from pre-test data. CARE represents seven evidence patterns, estimates the utility of 24 candidate interventions, and uses a risk gate to select context refinement. Across GIFT-Eval and six frozen TSFMs, CARE achieves consistent multi-metric improvements for all forecasters while revealing structured, forecaster-dependent intervention utility. These results suggest that treating context selection as a decision variable is a promising direction for evidence-aware inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.