TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Physiological Time Series
Abstract
Time Series Language Models (TSLMs) promise natural-language question answering over real-world temporal data, but their ability to retrieve and relate events in long time-series remains largely untested. We introduce TS-Haystack, a needle-in-a-haystack benchmark of ten event-grounded question-answering tasks over long physiological recordings, with contexts from 100 seconds to 9 hours, spanning direct retrieval, temporal understanding, multi-step retrieval, and contextual anomaly detection. Using this benchmark, we find that existing TSLMs exhibit performance degradation on long-context tasks: accuracy declines with context length, models that directly tokenize time-series without compression run out of memory at every context length, and other TSLMs collapse toward near-zero accuracy in interval retrieval tasks when increasing time-series length. These findings align with existing literature on text and multi-modal long context retrieval. We further contextualize TSLM performances with an oracle baseline and signal-blind baseline, and isolate their failure modes with an Agentic Retrieval for Time Series (ARTS) approach, a Large Language Model (LLM) agent with signal-specific classifiers as tools. Our findings uncover new failure modes in TSLMs at long context and show potential for architectures that further decouple local signal perception from global retrieval. Our code and datasets will be made publicly available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.