Forecast or Recall? Attributing Training Data Memorization in Time Series Foundation Models with TimeCanary
Abstract
Time series foundation models are evaluated on public benchmarks, yet their pretraining data are seldom documented at the series level. Low forecast error can thus reflect generalization or recall, leaving data owners unable to tell whether a model learned from their series. We find that a post-training audit depends on the baseline forecaster and lacks enough valid comparison pairs for a reliable conclusion. We introduce **TimeCanary**, which marks an owner's original series before release with a sparse Fourier signature scaled to each series. The mark extends across training windows. For each marked span, a recorded coin flip chooses one of two candidate keys for the mark; the other key provides a comparison. We score both keyed windows by model loss and use a matched model trained with reversed key roles to cancel pre-existing key differences. The test also requires reference models and diagnostic checks. To our knowledge, this is the first effective method shown to attribute use of marked original time series in a controlled audit. The paired score met its support rule across diverse models, including when only one quarter of eligible host spans were marked.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.