acceptodds
Under review as a conference paper at ICLR 2027

Byte-Pair Patching: Controllable Dynamic Tokenization for Time-Series Foundation Models

Abstract

Compression of input data, usually handled at the tokenization stage, is crucial for scaling time-series foundation models without incurring prohibitive computational costs. While static patch-based compression matches the performance of expensive point-wise tokenization, recent works have shown the scope for adaptive techniques that vary the patch size based on statistical properties of input data like entropy or frequency. To this end, inspired by Byte-Pair Encoding (BPE) for language, authors in Götz et al. (2025) expand an initial vocabulary of discretized numerical values by fusing them into motifs based on the frequency of occurrence in the training data. This method presents promising improvements in performance but only offers modest improvements to compression rates. In this paper, we introduce Byte-Pair Patching (BPP), a novel tokenizer that overcomes this limitation by applying BPE on fixed-size patches rather than point-wise observations, thus guaranteeing higher compression rates by design. Further, BPP also enables test-time control over compression, allowing a single pre-trained checkpoint to operate at different levels of input compression. Across GIFT-Eval, and 20 commonly used datasets, BPP provides up to increased compression versus point-wise tokenization and remains competitive relative to vanilla BPE and fixed-size patching baselines while improving inference latency by up to compared to the latter.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.