Bridging the Modality Gap between Time Series and Language for Cognitive Reasoning
Abstract
Time series analysis plays a central role in understanding dynamic systems and supporting forecasting, early warning, and decision-making. Recent studies have begun integrating time series with large language models (LLMs), yet most existing tasks remain focused on fixed-output prediction or state modeling, leaving the language understanding and reasoning capabilities of LLMs underexplored. Meanwhile, textualization-based approaches often fragment compact temporal structures and introduce substantial context overhead, while existing benchmarks provide limited reasoning depth and semantic diversity. To address these limitations, we introduce ***TSCognition***, a multimodal benchmark for comprehensive time series reasoning. It integrates real-world time series and textual information from ***15*** public sources and contains approximately ***41K*** question-answer pairs spanning five reasoning tasks: **Decoding**, **Grounding**, **Inferring**, **Extrapolating**, and **Acting**. Based on this benchmark, we further propose ***TSAlign***, a unified framework that directly aligns time series representations with the language representation space, enabling efficient and structured interaction between temporal signals and LLMs. Extensive experiments show that TSAlign consistently outperforms existing LLM-, VLM-, and time series QA-based methods on both TSCognition and the public TimerBed benchmark, while substantially reducing computational overhead. These results highlight representation alignment as an effective direction for scalable time series–language reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.