From General Experts to Channel Specialists: Recasting Mixture-of-Experts for Multivariate Time Series Forecasting
Abstract
Channels in a multivariate time series can favor different parameter updates. An expert's training channels therefore shape the competence available for forecast combination. We introduce cMoE, a plug-in mixture-of-experts mechanism that uses historical channel-expert forecasting errors to guide both training and prediction. Every expert evaluates every channel; an exponential-moving-average Scoreboard tracks channel-expert losses and determines the responsibility matrix. Unlike prediction-only weighting, this matrix both combines forecasts and defines each expert's channel-weighted training objective. Our analysis characterizes how responsibility changes expert updates and bounds future weighted expert-loss advantage over equal-expert assignment under exchangeability. Paired interventions on linear, frequency, and attention backbones show that responsibility-weighted training develops complementary experts and increases the average forecasting benefit of the same historical channel weights. Their value thus depends on the competence developed through training, not only on their use at prediction time. Across 180 settings with five backbones and nine multivariate datasets, cMoE improves every backbone and dataset on average, reducing MSE by 4.92% and MAE by 3.20%. Together, these findings support using outcome history to develop channel specialists as well as combine their forecasts. Our code is available at https://anonymous.4open.science/r/cMoE.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.