acceptodds
Under review as a conference paper at ICLR 2027

MoHETS: Long-term Time Series Forecasting with Mixture-of-Heterogeneous-Experts

Abstract

Real-world multivariate time series exhibit heterogeneous temporal behavior, including recurring structure, local variation, and non-stationary regimes, making long-horizon forecasting challenging. Sparse Mixture-of-Experts (MoE) approaches improve scalability and specialization, but they typically rely on homogeneous MLP experts that poorly capture the diverse temporal dynamics of time series data. We introduce MoHETS, an encoder-only Transformer that integrates sparse Mixture-of-Heterogeneous-Experts (MoHE) layers. MoHE combines an always-active shared depthwise-convolution path for sequence-level continuity with routed Fourier-based experts applied independently to latent patch tokens. The resulting layer is heterogeneous along three axes: operator class, receptive field, and activation role. MoHETS further improves robustness to non-stationary dynamics by incorporating exogenous information through a cross-attention pathway. We also replace parameter-heavy linear projection heads with a lightweight convolutional decoder, thus improving parameter efficiency and training stability. Finally, a single trained model autoregressively produces forecasts over arbitrary horizons. We evaluated MoHETS on seven established multivariate benchmarks across multiple horizons and consistently achieved superior performance, indicating that structural heterogeneity is important for long-term forecasting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.