acceptodds
Under review as a conference paper at ICLR 2027

Tabby-Pretrain: A Recipe for Continuable Pretraining of Time Series Foundation Models

Abstract

Pretraining a time series foundation model is costly, making it desirable for a training run to remain effective as it is extended. In this paper, we develop a pretraining recipe that supports continuable forecasting improvements across successive pretraining stages through the coordinated design of optimization, supervision, and data, and instantiate it in Tabby-Pretrain, an encoder-only Transformer-based TSFM. The resulting recipe combines a progressive convergence schedule (PCS), which repeatedly produces annealed checkpoints; deep quantile supervision (DQS), which directly supervises intermediate layers; and CauKerV2, an online generator that composes diverse temporal dynamics through randomly sampled structural causal models. Controlled experiments with the 40M-parameter Tabby-Small show that replacing PCS, removing DQS, or replacing CauKerV2 with additional real data degrades forecasting at the same training budget. We then apply the recipe to a 146M-parameter model. Its reported training trajectory improves across three stable–decay stages, and the final pretrained backbone is competitive in forecasting while supporting frozen representation transfer to classification and anomaly detection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.