Context in Time Series: Benchmarking and Aligning Time-Series Language Models
Abstract
Existing benchmarks for time-series language models can often be solved using either text or time-series alone, making it difficult to assess genuine cross-modal reasoning. In this paper, we introduce the Context in Time-Series Benchmark (CiT-Bench), which comprises forecasting, imputation, and anomaly detection tasks designed such that textual context is necessary for accurate prediction. The text states a future event or constraint, describes an event inside a missing block, or singles out the true anomaly among distractors. To facilitate model development on CiT-Bench, we additionally construct a language–time-series alignment dataset for training. Using this dataset, we show that improved understanding of both textual and time-series inputs can be achieved by training only the embeddings of a small vocabulary of time-series words, without fully fine-tuning the language model. We provide a unified framework for evaluating and improving cross-modal reasoning between language and time-series. Our benchmark is available at bluehttps://huggingface.co/anonymized.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.