VQ-LLMTS: Vector Quantization-based Large Language Models for Time Series Forecasting and Understanding
Abstract
Large Language Models (LLMs) excel at reasoning but struggle with time-series data due to the modality gap between discrete text tokens and continuous numerical signals. In this paper, we propose VQ-LLMTS, a multimodal framework that addresses this gap by discretizing continuous time series into a specialized temporal vocabulary. Unlike many alignment methods that directly bridge text and time-series embeddings, we employ Residual Vector Quantization (RVQ) to decompose complex temporal patterns into a hierarchical, discrete latent space. This expands the LLM's vocabulary with learned temporal tokens, enabling unified cross-modal reasoning while preserving the original backbone and single-head autoregressive architecture. To capture diverse temporal patterns, we develop a generalized hybrid data generator for RVQ codebook training. We then align the temporal vocabulary with the language model using synthetic examples for general capabilities and time-series understanding, together with real-world univariate time series. VQ-LLMTS achieves strong time-series understanding performance and competitive short-horizon forecasting performance using qualitative descriptions derived from the target windows. On the CiK benchmark, VQ-LLMTS also achieves the lowest context-conditioned RCRPS among the compared methods. These results support the effectiveness of an explicit temporal vocabulary for LLM-based time-series understanding and forecasting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.