Understanding and Improving Zero-Shot Forecasting through Attention-Head Functions in Time-Series Foundation Models
Abstract
Transformer-based Time-Series Foundation Models (T²SFMs) enable zero-shot forecasting across diverse domains by inferring temporal characteristics from unseen input contexts. However, how such capability is internally organized and utilized for forecasting remains poorly understood. Since self-attention performs input-dependent computation, attention heads provide a natural unit for investigating this question. We therefore introduce FBC, a Frequency-feature Based Causal framework for characterizing attention-head functions, revealing that individual heads are multifunctional yet functionally redundant, their functions exhibit both cross-domain stability and variation, and scaling functional heads causally alters forecasts. Together, these findings reveal how attention-head functions are organized and support zero-shot forecasting. Building on these, we propose SHIFT, a two-stage Soft-gated Head-wise Intervention Framework for improving T²SFMs' zero-shot forecasting capability. In calibration stage, SHIFT's soft-gated head-wise calibration module learns intervention vectors across observed domains and constructs a source-domain centroid bank. In inference stage, centroid-based retrieval module identifies a compatible source domain for unseen targets, and transfers its calibrated intervention vectors to modulate attention-head outputs, thereby enhancing zero-shot forecasting. Experiments on six T²SFMs across seven datasets demonstrate consistent zero-shot improvements, with average MSE and MASE reductions of up to 14.5% and 65.8%, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.