Time Series Compression with Foundation Models
Abstract
Time-series volumes continue to grow across domains, making efficient lossless compression essential for storage and transmission. Classical codecs and specialized numerical compressors often miss nonlinear or long-range temporal structure, while neural compressors that learn such structure typically demand per-dataset training and extra model storage. Large pre-trained foundation models offer strong sequence modeling with zero-shot transfer, suggesting a route to compression without task-specific fine-tuning. This paper connects next-token and next-value prediction to encoding-based compression and presents a unified zero-shot framework for numerical time series: LLMs drive arithmetic coding from predictive token distributions, and LTSMs drive residual coding from point forecasts. Lightweight numeric tokenization schemes are further studied without modifying frozen tokenizers. Extensive experiments show that publicly available foundation models can be applied directly; in particular, pre-trained LLMs attain higher compression ratios than the compared general-purpose, numerical, and LTSM-based compressors on most datasets, under a shared public-checkpoint accounting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.