acceptodds
Under review as a conference paper at ICLR 2027

Latency-Aware Time-Series Forecasting with Knowledge Distillation

Abstract

Time-series forecasters are scored offline as if inference took no time. On a stream sampled every Δt, the forecast of a call that takes T is first used d = ⌈T/Δt⌉ ticks later, when its first d steps describe instants already observed. We formalize latency-aware forecasting: a configuration issues its own stale prefix plus the H points the consumer needs and is scored on the remaining usable window. This deployment error factorizes into a device-independent error curve and a device-dependent prefix, so candidates are trained once and re-scored per device. Its optimum can be interior, since larger models lower per-step error but push the usable window further out. On a 1 kHz ECG stream and one x86 core, the usable-window error selects an MLP student a quarter the size of the one that nominal-window accuracy picks. For a selected configuration, reweighting the loss toward the usable window adds no signal to per-step read-outs and little to shared encoders, so TRIAD changes the target instead. It invokes the teacher d steps after the student, so that the teacher's forecast covers exactly the usable window, which helps when the teacher's bias grows with lead time, and adds frequency and correlation terms. With a ModernTCN teacher, TRIAD lowers the error of the selected student by about 3% on ECG and 6% on ETTh2 replayed at 1 ms, in every latency setting. On a 1 kHz gait stream it gains 3 to 12% in two latency settings and loses 0.6% in the third. Removing its frequency or correlation term raises the error by at most 0.8%. Code is available at https://anonymous.4open.science/r/TRIAD-8B7D.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.