Investigating the effects of patch size in Time Series Foundation Models
Abstract
Time series analysis is a core task across science, industry, and decision-making, and transformer-based foundation models have recently emerged as a promising unified approach. A critical step in designing transformer-based Time Series Foundation Models (TSFMs) is tokenization. Most existing TSFMs use patch-based tokenization, making the patch size a central design choice that controls temporal resolution, token sequence length, and computational cost. In this work, we systematically study the role of patch size and architecture choice in time series foundation models, two decisions that are typically studied independently but, as we show, are deeply coupled. We compare encoder-only and decoder-only forecasting models trained on the same data mixture and with the same optimization budget, varying the patch size within each model class. The two architectures respond differently: across 97 GIFT-Eval tasks, encoder-only models improve as patch sizes decrease, whereas decoder-only models perform best at an intermediate patch size, reflecting a dataset-level preference that shifts towards larger patches as the forecasting horizon grows. On UCR-91 classification with frozen representations, smaller patches again help the encoder, while the decoder is hardly affected. Encoder-only models also benefit more from longer context windows, though gains saturate beyond 2048 context length. The two architectures occupy complementary positions on the accuracy–compute trade-off: decoder-only models are more efficient, while encoder-only models achieve higher accuracy. These findings suggest that patch size and architecture should not be chosen in isolation, but should both follow from the target forecasting horizon and deployment constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.