Future-Guided On-Policy Self-Distillation for Time-Series Forecasting
Abstract
Time-series forecasting has a training asymmetry: each training example eventually reveals the future, but a deployed model sees only history. That future contains both history-predictable structure and innovations. Standard pointwise supervision uses the realised values as targets; can their ordered structure also guide learning without entering the deployed forecast? FoldOPSD turns this training-time context into on-policy self-distillation. A teacher view of the current forecaster receives summaries of observed future segments, while a history-only student uses a learned constant correction. Their shared forecasting body lets teacher supervision update the model that will be deployed; detached residual matching and an EMA student add teaching targets. Because both views share the forecasting body, the answer changes the training update rather than the deployed input. Analysis identifies when ideal answer conditioning reduces gradient noise and how answer hints reweight residual directions. At a fixed DLinear/ETTm2 checkpoint, the joint shared-body gradient lowers reference MSE by 28.3% on training and 35.0% on validation, while across-batch variance falls by 33.4% and 33.9%, respectively. With a compatible native output bias, the student correction folds exactly into that bias, adding no inference operations. Across four datasets and five seeds, FoldOPSD lowers mean test MSE in all 24 DLinear and iTransformer settings, by up to 10.85%. True future hints outperform noise and dataset-level summaries in all 16 matched ETT settings with the basic OPSD objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.