DaTS-Mix: Selecting Pretraining Mixtures for Time-Series Foundation Models Using Domain and Temporal Statistics
Abstract
Domain-aware pretraining-mixture selection has been widely studied for large language models, but its use in time-series foundation model (TSFM) pretraining is less well established. We study this problem and emphasize temporal statistics as a complement to domain information when selecting source proportions. We propose DATS-MIX, which represents each mixture using source proportions, domain shares, and the first and second raw moments of 14 temporal statistics computed under the training sampler. We condition a TabPFN response predictor on outcomes from 128 runs of a 7.07M-parameter forecasting proxy and use it to select a mixture for larger-model pretraining. On later evaluation windows from the same 28 sources in seven domains, the selected mixture reduces geometric-mean forecasting loss by 10.85% relative to size-proportional sampling and by 5.57% relative to a matched selector without temporal statistics, under the same 20k-update budget. The mixture selected with both moment blocks achieves the lowest mean native loss and weighted quantile loss among the evaluated methods; a Mean-only ablation separates the incremental role of the second-moment block. These results support proxy-guided TSFM data-mixture design and highlight the value of temporal statistics beyond source and domain proportions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.