One Patch, Three Roles: What Is Actually Coupled in Autoregressive Time-Series Forecasting?
Abstract
Patch-based autoregressive time-series forecasting often ties input representation, learned transition span, and per-call output length to one patch length. We investigate whether these roles can be adjusted separately. For encoding, we explore building patches from short temporal atoms. We find that grouping preferences vary across data sets and model widths, but performance is more sensitive to width than to the number of atoms per patch on the evaluated grid. For decoding, we ask whether models can predict more subsequent patches per call. Lightweight parallel exits fit a parent’s recursive trajectory substantially more easily than the observed future, motivating Autoregressive Trajectory Distillation (ATD). ATD compiles the frozen parent’s rollout into selectable multi-patch execution. ATD-8 achieves a 5.54× end-to-end speedup with stable quality across widths in a paired four-data-set comparison. However, imitation can preserve the parent’s accumulated forecast error. We therefore ask whether this error can be mitigated independently of execution width. We find a correctable residual projection along a train-selected periodic history direction, motivating Spectrum Tangent. This correction adds no neural parameters or Transformer calls. At horizon 720, it reduces mean squared error (MSE) and mean absolute error (MAE) by 2.54% and 2.33% across seven data sets and two execution widths, retaining a 3.24× speedup over recursive inference. Experiments on three public autoregressive parents and four data sets demonstrate both methods’ transferability, with Tangent providing further long-horizon improvements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.