Spatio-Temporal Decomposition: A Plug-in Interpretable Front End for Multivariate Time Series Forecasting
Abstract
Forecasting models for spatio-temporal series usually either ignore cross-location structure or encode it with a graph that is fixed in advance and hard to read. We propose Spatio-Temporal Decomposition (STD), a module that sits in front of any forecasting backbone and splits each observation into four named components, namely local and global in space and high and low frequency in time. The split uses only means over time and over locations plus two sets of learned offsets, one per location and one per time step, so it introduces no additional hyperparameters beyond those of the backbone. We show that the per-location offset converges to the correction a piecewise-stationary trend requires, and that the per-time offset acts as an implicit set of spatial weights, so the plain average over locations becomes a learned neighborhood without a graph. Across TCN, Transformer and WaveNet backbones on BikeNYC, SKT, and Traffic, STD lowers MAE in 26 of 27 backbone, dataset and horizon combinations, cuts MAE summed over the three horizons by 19 to 29% on BikeNYC, and lifts graph-free backbones above MTGNN and StemGNN. On BikeNYC the global high-frequency part shows two commuting peaks on weekdays and one on weekends, indicating that STD surfaces readable structure that graph-based models encode only implicitly. Ablation experiments consistently show that removing spatial aggregation increases mean RMSE, highlighting the contribution of spatial information to forecasting accuracy. Robustness evaluations on Capital further show that STD maintains stable performance under Gaussian noise and achieves the lowest MAE and RMSE among competing models at 20% missingness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.