acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Autoregressive Time Series Pretraining with Next--Token Prediction

Abstract

Autoregressive pretraining is a central paradigm for building foundation models, yet its effectiveness for time series foundation models (TSFMs) remains controversial. Conventional next-token prediction (NTP) provides only local supervision and suffers from error accumulation. Existing Next--Token Prediction (NNTP) approaches that predict multiple future tokens per step alleviate these issues to some extent, but introduce a fixed-horizon bias that may hinder generalization across diverse forecasting horizons. To address these limitations, we reformulate autoregressive time series pretraining as Next--Token Prediction (NxTP), which dynamically samples the prediction horizon for each token based on the available context. By enabling models to directly predict sufficiently long future sequences within a single decoding step, NxTP substantially reduces the need for autoregressive rollouts while improving generalization across forecasting horizons. Extensive experiments on two well-established benchmarks demonstrate that NxTP consistently improves forecasting performance and horizon generalization while substantially accelerating inference in long-term forecasting scenarios.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.