acceptodds
Under review as a conference paper at ICLR 2027

AgenTS: Harnessing Runtime Tool Orchestration for Time Series Reasoning

Abstract

Time series reasoning involves understanding temporal patterns and reasoning over them to answer complex questions across domains such as healthcare, energy, and climate science. Large language models often struggle to abstract relevant numerical structure from long and noisy time series. Analytical tools can support this abstraction, but their heterogeneous capabilities and output structures require coordination according to the question and input data. To address these challenges, we propose AgenTS, a runtime harness that coordinates task analysis, tool configuration, execution, and evidence preparation for time series reasoning. The core idea is to disassemble inference into tool-based quantitative abstraction and evidence-based reasoning, making the abstraction process an explicit target of runtime control. Specifically, a Profiler constructs a per-instance runtime state that jointly characterizes task relevance and input series properties. An Orchestrator uses this state to construct and select ordered tool configurations under a utility criterion balancing task coverage, question relevance, task–data capability match, and computational cost. Procedural memory supports configuration reuse across similar instances, reducing repeated configuration generation. A Prober executes the selected configurations and aggregates their heterogeneous outputs into structured evidence for subsequent model reasoning, with final answers obtained through multi-sample voting. Experiments on two benchmarks covering 14 datasets and 10 reasoning tasks demonstrate strong performance, with an average improvement of more than 30% across tasks while using 33% fewer tools than the strongest baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.