acceptodds
Under review as a conference paper at ICLR 2027

Revisiting Recurrent Models for Long-Term Time Series Forecasting through Representation Accessibility

Abstract

Transformer-based architectures have become prominent in long-term time series forecasting (LTSF), and their empirical advantages are often attributed to self-attention. This paper revisits that attribution by identifying a less visible architectural confounder in common LTSF comparisons: representation accessibility. In many direct multi-horizon recurrent baselines, the prediction module is conditioned on a compressed recurrent summary, typically the final encoder state, or on a decoder initialized from that state. In contrast, patch-based Transformer forecasters such as PatchTST expose representations from multiple temporal positions directly to a non-autoregressive prediction head. This difference changes not only the sequence modeling mechanism, but also which historical representations are accessible at prediction time. To isolate this factor, we introduce a Full-State Readout (FSR) framework built from standard recurrent encoders. FSR keeps key ingredients of strong LTSF forecasters fixed, including patching, channel independence, normalization, and a Flatten–Linear prediction head, while varying only the number of recurrent hidden states exposed to the head. Across six LTSF benchmarks, FSR with a GRU backbone matches or outperforms representative Transformer baselines. Expanding representation accessibility generally reduces forecasting errors for GRU, and full-state readout also improves vanilla RNN and LSTM encoders. These results show that the recurrent–Transformer gap in LTSF cannot be attributed solely to recurrence versus self-attention. More broadly, they suggest that final-state bottlenecks can systematically underestimate recurrent models, and that representation accessibility should be controlled when comparing sequence modeling architectures for forecasting.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.