When Representation Change Misleads: Evaluating Encoder Adaptation in Time-Series Foundation Models
Abstract
Fine-tuning a time-series foundation model reshapes its internal representations, and that reshaping is routinely read as catastrophic forgetting to be prevented. Preventing it presupposes something worth preserving. **Centered kernel alignment (CKA) tells us that the representation changed; a fine-tune-versus-frozen-encoder intervention tells us whether that change was valuable. Across the cells we evaluate, the former is not an actionable substitute for the latter.** Screening **32 cells** (three backbones: Moirai-S/B/L, Chronos-T5, TimesFM-2.5; six datasets; two horizons) against a lookback-96 linear regression *fit on the target data* and scored on windows disjoint from every outcome we report, only **7 of the 31** scored cells show the pre-trained model beating that baseline, and all seven are Moirai on two of the six datasets. On those seven, full fine-tuning *improves* MSE in five. **No cell meets our prespecified criterion for harmful encoder adaptation**: under multiplicity-adjusted paired intervals, the 31 cells split **2 freezing-better, 4 adaptation-better, 25 inconclusive**, and both cells where freezing decisively wins are cells our own screen rejects; a pre-registered seed top-up adds a third it *admits*, and the pre-registered unanimity form returns zero throughout. The outcome nonetheless spans to against zero-shot, and under an analysis fixed before it was run we find **no evidence** that CKA improves out-of-sample prediction of it beyond backbone identity: across 15 leave-one-cluster-out folds it *lowers* from to . In sample CKA **does not provide a reliable, actionable ordering** either: within Moirai, the only backbone with enough cells to ask, (clustered CI ), and acting on it as prescribed (freeze where drift is largest) is worth pp against always adapting. Two measurement choices decide this: the frozen-encoder comparison must be scored on **held-out** windows (4 of 31 cells otherwise reverse sign, every one of them flattering the frozen encoder, and two of the four are cells that pass the screen), and the baseline must be **fitted**, since an unfitted per-window trend extrapolation that a constant predictor beats puts 16 of 21 Moirai cells on the wrong side of the screen.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.