LEXIS: A Symbolic Vocabulary for Time Series Understanding in Pretrained Language Models
Abstract
Pretrained large language models (LLMs) exhibit strong language understanding and reasoning capabilities, but applying these capabilities to time series understanding requires an interface that exposes temporal structure in a form legible to the model. Existing interfaces represent a series either as decimal strings or as embeddings produced by a series encoder. Decimal strings preserve the precise values but obscure their temporal relationships, while encoder embeddings are compact and information rich but lack semantics known to an LLM. We propose , an interface that expresses a time series as language by extending the vocabulary of a pretrained LLM. A small set of new symbols describes the structure of a series, each symbol denoting one known component of a curve, and their embeddings are derived from the geometric relationships among the symbols. Across three time series understanding benchmarks, LEXIS improves a Qwen3-4B backbone by 31.09%, 13.40%, and 4.45% and surpasses evaluated general-purpose LLMs and time series language models on most tasks. Controlled comparisons show that LEXIS outperforms the decimal and encoder embedding representations under the same training recipe. Ablations show that the model effectively derives temporal properties from the LEXIS representation and benefits consistently from the symbol geometry. These results establish LEXIS as an effective and principled interface between time series and pretrained LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.