acceptodds
Under review as a conference paper at ICLR 2027

One Patch, Three Roles: What Is Actually Coupled in Autoregressive Time-Series Forecasting?

Abstract

Patch-based autoregressive time-series forecasting often ties input representation, learned transition span, and per-call output length to one patch length. We investigate whether these roles can be adjusted separately. For encoding, we explore building patches from short temporal atoms. We find that grouping preferences vary across data sets and model widths, but performance is more sensitive to width than to the number of atoms per patch on the evaluated grid. For decoding, we ask whether models can predict more subsequent patches per call. Lightweight parallel exits fit a parent’s recursive trajectory substantially more easily than the observed future, motivating Autoregressive Trajectory Distillation (ATD). ATD compiles the frozen parent’s rollout into selectable multi-patch execution. ATD-8 achieves a 5.54× end-to-end speedup with stable quality across widths in a paired four-data-set comparison. However, imitation can preserve the parent’s accumulated forecast error. We therefore ask whether this error can be mitigated independently of execution width. We find a correctable residual projection along a train-selected periodic history direction, motivating Spectrum Tangent. This correction adds no neural parameters or Transformer calls. At horizon 720, it reduces mean squared error (MSE) and mean absolute error (MAE) by 2.54% and 2.33% across seven data sets and two execution widths, retaining a 3.24× speedup over recursive inference. Experiments on three public autoregressive parents and four data sets demonstrate both methods’ transferability, with Tangent providing further long-horizon improvements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.