acceptodds
Under review as a conference paper at ICLR 2027

Tstar: Temporal Structure-Aware Training Data Construction for Time-Series Models

Abstract

Time-series foundation models increasingly rely on large, heterogeneous pretraining datasets, yet how to construct these datasets has received far less systematic study than in natural language processing and computer vision. Dataset scale and source diversity alone do not determine how training exposure is allocated across temporal structures, so even a diverse corpus can leave some structures rarely seen. To address this, we introduce Tstar, a temporal structure-aware method that organizes training windows from heterogeneous datasets into an interpretable atlas of temporal structures. The atlas guides real-window selection and sampling, with optional targeted synthesis, to increase exposure to underrepresented patterns. We pretrain TinyTimeMixer and MOMENT-341M from random initialization under fixed budgets. For TinyTimeMixer, coverage-guided selection reduces the six-seed means of four TIME and FEV error aggregates by 23.41%–41.81% relative to matched uniform selection, with every paired seed improving. For MOMENT-341M, real-window rebalancing alone reduces MASE by 3.43% on TIME and 7.67% on FEV relative to uniform sampling, with both 95% bootstrap intervals excluding zero. A supporting study on six conventional architectures shows lower average error.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.