acceptodds
Under review as a conference paper at ICLR 2027

Reshaping Tool Calling Policies of Time Series Agents via Reinforced Evaluation

Abstract

Time series reasoning, which involves understanding temporal patterns to answer complex questions, plays a crucial role in a variety of domains such as healthcare, energy, and climate. Modern approaches have tried using large language model (LLM) based agents with analytical tools to reason over time series. However, these methods would always execute tool calls directly without evaluation, which inevitably leads to many low-value tools being executed and uninformative results being generated, thus hindering both the reasoning efficiency and accuracy. One seemingly workable solution inspired by LLM and agent studies is to introduce post-training to guide agents for wiser tool calling. However, tuning these massive parameters of LLM-based agents further imposes additional computational burden. To address these problems, in this paper, we explore a novel direction of evaluating time series reasoning without post-training base agents. The core idea is to bring an evaluation reward model to assess each tool call using accumulated evidence, whose evaluation can further illuminate exploration of time series agents and reshape their action strategies. Naturally, we reformulate the iterative evaluation and action-reshaping process into a Markov decision problem that maximizes the expected reward composed of tool costs and task performance. To optimize the objective, we propose a novel time series agentic framework, TimesEval, which uses value-based reinforcement learning to learn evaluation policies for reshaping agent actions. To estimate rewards for each evaluation step, we propose episode relative advantage learning, which measures the expected return difference among decisions as the advantage for fast convergence of action reshaping. Moreover, to approach the global optimal policy, we utilize evolving policy optimization, which approximates policy iteration with on-policy state sampling to progressively improve evaluation reward model. Notably, the key difference is TimesEval trains a lightweight evaluator through reinforcement learning to reshape agent actions instead of post-training the agent itself. Finally, we conduct extensive experiments on 15 datasets and across four time series tasks, demonstrating that TimesEval consistently achieves state-of-the-art performance, with an average improvement of more than 19% over the strongest baseline and 36% lower tool cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.