acceptodds
Under review as a conference paper at ICLR 2027

A Forecasting Odyssey: Charting the Billion-Scale Landscape of Time-Series Forecasting

Abstract

Time-series forecasting—using historical patterns to predict future outcomes—has gained significant attention due to its broad practical impact nowadays. Despite decades of methodological advances and benchmark development, existing evaluation frameworks still exhibit critical limitations: (i) many studies rely on homogeneous or imbalanced datasets with limited cross-domain diversity and quality assurance; (ii) model coverage is often limited, overlooking simple yet effective baselines; and (iii) insufficient statistical and domain-specific analyses lead to superficial conclusions. To address these limitations, we introduce TimeFCST, a billion-scale benchmark with systematic quality assurance that evaluates 100 forecasting methods, roughly 5 more than prior studies, across statistical, machine learning, deep learning, and foundation models. Beyond aggregate leaderboards, we provide, to our knowledge for the first time, rigorous statistical validation of comparative findings over large-scale experiments and a failure-profiling framework spanning six error patterns. Together, these answer three questions: What works, when does it work, and how does it fail? Our evaluation reveals: (i) modern forecasting models lead aggregate rankings, yet statistically significant gains over well-tuned simple classical baselines emerge only among leading foundation models especially on UTS; (ii) model effectiveness varies across domains and settings, with simple classical methods matching or even outperforming modern models in domains such as nature, retail and economics; (iii) failure profiles reveal distinct error patterns hidden by aggregate metrics: deep learning and foundation models exhibit more phase and localized errors, with foundation models showing stronger underprediction bias; and (iv) additionally, we identify saturated datasets that obscure methodological progress and show that broad, diverse benchmark coverage reveals otherwise hidden, statistically significant differences among models. In summary, TimeFCST charts the evolving landscape of time-series forecasting, offering insights into recent advances and directions for future research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.