Adaptive Quantile Response-Based Knowledge Distillation for Time-Series Foundation Models
Abstract
Knowledge distillation has been widely used to transfer knowledge from large foundation models to compact students, yet remains underexplored for time-series foundation models (TSFMs), particularly in multi-teacher probabilistic forecasting. To address this gap, we propose a multi-teacher response-based knowledge distillation framework that consolidates complementary forecasting knowledge from multiple pretrained TSFMs into a single compact student TSFM. Our method introduces Order-Huber Relative KD (OH-RKD) for teacher-guided cross-quantile distillation and Sample-wise Relative Adaptive Weighting (SRAW), which adaptively adjusts the distillation strength of the selected teacher according to its forecasting quality relative to the student. Extensive experiments on FEV-Bench and GIFT-Eval show that our method reduces SQL on FEV-Bench by 9.03% and WQL on GIFT-Eval by 10.96% relative to the matched GT-only student, demonstrating the effectiveness of multi-teacher distillation for TSFMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.