acceptodds
Under review as a conference paper at ICLR 2027

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models

Abstract

Time series foundation models (TSFMs) are increasingly pretrained on large corpora, raising concerns that evaluation datasets may have been exposed during pretraining and thus yield overly optimistic performance estimates. Auditing such contamination is challenging in time series because signals are continuous and heterogeneous, and often lack corpus documentation. We formalize pretraining source-exposure auditing for TSFMs and propose TSFMAudit, a method based on probe adaptation dynamics. Our key hypothesis is that prior exposure can support efficient adaptation: a short fine-tuning probe may achieve faster loss reduction with smaller backbone movement. Candidate-to-reference differences and ratios compare this behavior across heterogeneous forecasters. We evaluate TSFMAudit on six TSFMs and 137 datasets using documented training-source evidence, with a separate evaluation on 50 TIME datasets. Against documentation-derived labels, combining reference-relative and candidate features yields candidate-averaged MCC 0.317, compared with 0.050 for the strongest tested scalar baseline and 0.169 for candidate-only dynamics. We further validate TSFMAUDIT under ground-truth membership by pretraining Chronos-T5-mini and Moirai-1.0-R-Small from random initialization on controlled dataset manifests, obtaining MCCs of 0.356 and 0.618, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.