TempoDraft: Efficient Long-Horizon Rolling Forecasting
Abstract
Recent time-series foundation models (TSFMs) achieve strong long-horizon forecasting performance, but their deployment for rolling forecasting remains costly. Their autoregressive decoding invokes the model sequentially for each forecast token. In autoregressive LLM generation, new tokens extend an unchanged prefix, allowing its KV cache to be reused across decoding steps. Rolling forecasting does not preserve this structure. Each new observation advances the fixed-length input window by one time step. This changes the input patches and their causal context, preventing direct reuse of the previous KV cache. However, the previous forecast still covers most timestamps in the new horizon. Thus, we introduce TempoDraft, a training-free method that recasts rolling forecasting as a refinement of the previous forecast under new observations. This forecast provides a temporal self-draft for the next inference step. TempoDraft first rebuilds the current context state and computes the first forecast token exactly as in autoregressive decoding. It then shifts the previous forecast to form an aligned draft and refines the remaining forecast tokens in parallel. A rolling forecast ensemble combines current and earlier forecasts of the same timestamps to stabilize predictions and improve accuracy. On Timer, extending the forecast horizon on hourly ETTh1 from approximately one month to one quarter increases the decoder speedup over autoregressive decoding from 2.3× to 6.6×. At long horizons, TempoDraft reduces MSE by 3.3% on ETTh1 and 6.6% on Electricity relative to the autoregressive baselines. Further experiments across nine benchmarks and multiple TSFM backbones support the method's generalizability and broad applicability
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.