ORCHESTRA: When the Best Teacher Is Not the Best Guide in Multi-Teacher Time Series Distillation
Abstract
Knowledge distillation (KD) can transfer the complementary strengths of time-series forecasting models to a compact student. However, temporal patterns differ across forecasting tasks, so the best-performing teacher may change with context, and even its guidance may be difficult for the current student to imitate. Choosing or combining multiple teachers does not by itself determine how strongly the student should learn each teacher’s temporal signals. To address these challenges, we propose ORCHESTRA, a multi-teacher distillation framework that separates learning into credit, trust, and execution. A reinforcement-learned Teacher-Credit Agent uses the Adaptive Credit Policy (ACP) to learn context-dependent teacher credit. The Multi-Scale Trust Estimator (MSTE) calibrates point, scale, and first-difference supervision based on student–teacher discrepancies, while the Sparse Execution Policy (SEP) converts learned credit into deterministic sparse distillation weights. Extensive experiments on seven forecasting datasets and four prediction lengths show that ORCHESTRA consistently improves the compact TSMixer student, achieves the best overall MSE and MAE among the evaluated direct forecasters, and retains efficient student-only deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.