acceptodds
Under review as a conference paper at ICLR 2027

A Verifier from the Past: Test-Time Compute for Time Series Foundation Models

Abstract

Test-time compute (TTC) improves language models by spending additional inference compute to generate, verify, and select among candidate outputs, but this recipe has not transferred directly to time series forecasting. A language model can verify a proof or vote across samples, but a forecaster must be causal; it cannot use future information to score a candidate. We propose Backtest-Gated Augmentation Averaging (BGAA), a TTC algorithm for time series forecasting that builds a causal verifier from the series’s own history, using rolling-origin backtests to score structured views of the context, discard those that underperform the untransformed context, and average the survivors’ predictive distributions. BGAA improves every tested TSFM–benchmark combination, reducing MASE by 0.5–1.1% on GIFT-Eval, 3.2–5.9% on fev-bench, and 1.9–5.1% on BOOM, where a 313M model with TTC outperforms the 1B and 2.5B models in its family. We further show that useful views have structured, model-dependent effects consistent with a prior-transport interpretation. On the 53 cross-fittable GIFT-Eval configs, the deployed policy gains only 0.4–1.5% CRPS, versus 1.4–4.1% when choosing the best view for each dataset configuration and 17–19% when choosing the best view separately for each test window. This gap suggests that a major remaining challenge is window-level view selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.