TRACS: Time-Series Rebalancing and Coverage-Aware Selection for Pretraining
Abstract
As time-series foundation models (TSFMs) mature, performance increasingly depends on the quality and composition of their pretraining data. Most existing data selection methods rely on training feedback or external supervision, which makes it difficult to determine the selected corpus entirely before pretraining. We present TRACS (Time-Series Rebalancing and Coverage-Aware Selection), a zero-feedback method for selecting time-series pretraining data without either source of supervision. TRACS uses GMM-based soft clustering to organize candidate series by temporal pattern, then assigns sampling quotas based on the coverage required within each cluster rather than on cluster size. The selected subset and its pattern coverage can therefore be determined before pretraining begins. Under the stated assumptions, we prove coverage of all reachable patterns, finite-step termination of the missing-pattern repair procedure, and a bounded representation error with respect to the rebalanced target distribution. We evaluate TRACS on MOMENT-1 Small, PatchTST, and Timer under the same 10B-token pretraining budget. Using only 25% or 50% of the candidate corpus, TRACS attains the lowest aggregate MASE and CRPS among the evaluated pretraining baselines on 20 stable zero-shot forecasting tasks from GIFT-Eval. Relative to pretraining on the full corpus under the same token budget, TRACS reduces MASE by 2.2%–4.0% and provides an explicit guarantee on pattern coverage that the baseline methods do not offer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.