From Forecasts to Anomaly Scores: Reusing Frozen Time Series Foundation Models
Abstract
Frozen forecasting time series foundation models (TSFMs), when used for anomaly detection based on one-step forecast error (conventional scoring), have been reported to perform comparably to a model-free baseline that predicts by averaging the last three observations. We investigate whether their existing multi-step forecasts contain information discarded by conventional scoring. We propose Forecast Disagreement Anomaly Scoring (FDAS), a post-hoc scoring layer that combines the one-step forecasting error with disagreement among forecasts for the same target time made at different forecast origins. FDAS smooths and normalises both signals using a reference segment presumed to contain normal observations, then takes their maximum as the anomaly score. FDAS reuses forecasts already produced by a TSFM, without fine-tuning it or using anomaly labels for the target series. On the full TSB-AD benchmarks, we evaluate FDAS on eight univariate and five multivariate TSFM configurations. Compared with conventional scoring, FDAS raises mean VUS-PR from 0.29–0.30 to 0.45–0.50 on the univariate benchmark and from 0.21–0.24 to 0.37–0.40 on the multivariate benchmark. These gains are statistically significant in every evaluated configuration. Ablations show that combining the smoothed forecasting-error and disagreement tracks outperforms either track alone. FDAS makes frozen forecasting TSFMs competitive with dedicated anomaly-detection TSFMs, indicating that their limitation lies largely in the scoring rule rather than the forecasts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.