What Does Language-Model Pretraining Contribute to Time-Series Understanding
Abstract
Our contribution is a controlled empirical attribution analysis of the role of language-model pretraining in time-series understanding across decision targets and output interfaces. Across ten base language models, we compare frozen pretrained backbones with architecture-matched random controls under matched inputs, feature extraction, and learned-readout training. In controlled forecast- ing, pretraining improves process-family routing in all ten models, with gains of 7.5-31.4 percentage points, but the effect reverses when supervision targets the future-optimal expert; a five-model expected-risk control preserves this reversal. Affine forecasting heads instead reduce error in nine of ten models. On four real-data domains, task-specific readouts improve correct updating after evidence edits by 14.80 points on average, while preservation of unchanged facts decreases by 7.59 points. Under native constrained decoding, models retain the pre-edit answer in 88 of 100 evaluations conditioned on an initially correct target. Readout controls show that the measured contribution depends on how temporal features are accessed. Together, the analysis shows that pretraining supplies reusable temporal features, but its benefit is decision- and interface-dependent rather than a uniform gain in time-series understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.