acceptodds
Under review as a conference paper at ICLR 2027

LEXIS: A Symbolic Vocabulary for Time Series Understanding in Pretrained Language Models

Abstract

Pretrained large language models (LLMs) exhibit strong language understanding and reasoning capabilities, but applying these capabilities to time series understanding requires an interface that exposes temporal structure in a form legible to the model. Existing interfaces represent a series either as decimal strings or as embeddings produced by a series encoder. Decimal strings preserve the precise values but obscure their temporal relationships, while encoder embeddings are compact and information rich but lack semantics known to an LLM. We propose , an interface that expresses a time series as language by extending the vocabulary of a pretrained LLM. A small set of new symbols describes the structure of a series, each symbol denoting one known component of a curve, and their embeddings are derived from the geometric relationships among the symbols. Across three time series understanding benchmarks, LEXIS improves a Qwen3-4B backbone by 31.09%, 13.40%, and 4.45% and surpasses evaluated general-purpose LLMs and time series language models on most tasks. Controlled comparisons show that LEXIS outperforms the decimal and encoder embedding representations under the same training recipe. Ablations show that the model effectively derives temporal properties from the LEXIS representation and benefits consistently from the symbol geometry. These results establish LEXIS as an effective and principled interface between time series and pretrained LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.