Revisiting Embeddings in Time Series Forecasting from Latent Space and Data Perspective
Abstract
Embeddings serve as the fundamental representation units in time series forecasting and impose critical inductive biases on model learning, substantially influencing forecasting performance. Despite their importance, existing studies on embeddings largely overlook the distinction between direct mapping and high-order interactions on embedding effectiveness. In this work, we systematically investigate the effectiveness of embeddings across diverse datasets, ranging from first-order mapping to high-order interactions. Through extensive large-scale hyperparameter exploration and evaluation-theoretic analysis, we find that embedding-based models often degenerate into simple linear mappings without high-order interactions. In contrast, under higher-order interactions, embeddings exhibit distinct sensitivities to specific data characteristics and demonstrate distinct behaviors under different encoder architectures, suggesting that embeddings cannot be studied independently of higher-order interaction mechanisms. Our forecasting comparison covers attention, MLP mixing, recurrent, state-space, frequency-based, and convolutional encoders. Dataset characteristics, spectral rank, and frozen linear probes connect view preferences to representation evolution. The spectral and probe analyses cover MLPMixer, Transformer, GRU, Mamba, and FFT. With fixed-budget probes, encoder outputs reduce readout error in 72 of 98 configurations, with gains varying by dataset and token view. A multiscale data-rank rule further selects embedding–encoder pairs before target-model training and recovers six source-derived reference choices under leave-one-dataset-out calibration. Together, these results connect view construction, representation evolution, and practical module selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.