Shifted-window Diffusion with Temporal-distance Attention for Self-Supervised Time-Series Forecasting
Abstract
Scaling drives modern machine learning, but its usual currency is parameters, data, and GPU memory. Holding the model and data fixed, we instead scale the supervision, namely the number of prediction targets per input. Existing self-supervised pretexts for time series reconstruct targets within a single window or span at most two, so none of them can scale this axis. We propose SDTA (Shifted-window Diffusion with Temporal-distance Attention), a self-supervised pretraining framework whose encoder is fine-tuned for downstream forecasting. SDTA encodes the input window once and denoises diffusion-noised target windows, each shifted a sampled distance into the future, so that one encoding must support prediction at many temporal distances. Temporal-distance attention injects each target's distance only into the attention keys, steering where the decoder attends rather than what it copies. Because the input is encoded once and the targets exist only in pretraining, raising multiplies supervision at linear activation memory, fits one consumer GPU, and adds no deployment cost. Experiments on twelve standard forecasting benchmarks spanning energy, weather, finance, and traffic compare eight self-supervised methods under identical encoder capacity, one fixed configuration, no per-dataset tuning, and five-seed averaging. We further scale each baseline's own supervision axis: most possess no memory-feasible counterpart, and the two that do gain nothing from scaling it, whereas SDTA improves consistently along this axis, with the gain widening beyond the standard horizons. SDTA also achieves the lowest average MSE and MAE among the eight methods, and ranks top-two on ten of the twelve benchmarks by MSE and nine by MAE.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.