MoHETS: Long-term Time Series Forecasting with Mixture-of-Heterogeneous-Experts
Abstract
Real-world multivariate time series exhibit heterogeneous temporal behavior, including recurring structure, local variation, and non-stationary regimes, making long-horizon forecasting challenging. Sparse Mixture-of-Experts (MoE) approaches improve scalability and specialization, but they typically rely on homogeneous MLP experts that poorly capture the diverse temporal dynamics of time series data. We introduce MoHETS, an encoder-only Transformer that integrates sparse Mixture-of-Heterogeneous-Experts (MoHE) layers. MoHE combines an always-active shared depthwise-convolution path for sequence-level continuity with routed Fourier-based experts applied independently to latent patch tokens. The resulting layer is heterogeneous along three axes: operator class, receptive field, and activation role. MoHETS further improves robustness to non-stationary dynamics by incorporating exogenous information through a cross-attention pathway. We also replace parameter-heavy linear projection heads with a lightweight convolutional decoder, thus improving parameter efficiency and training stability. Finally, a single trained model autoregressively produces forecasts over arbitrary horizons. We evaluated MoHETS on seven established multivariate benchmarks across multiple horizons and consistently achieved superior performance, indicating that structural heterogeneity is important for long-term forecasting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.