acceptodds
Under review as a conference paper at ICLR 2027

Can Language Models Program Better Time-Series Anomaly Detectors?

Abstract

Language-model agents for time-series anomaly detection typically inspect series directly or orchestrate existing detectors to identify and explain anomalies. We introduce AgentAD (Agent As Developer), in which the agent writes a standalone Python detector rather than serving as the detector itself. The resulting program runs without any language-model calls. We design an evidence-driven loop to help the agent develop detector programs that generalize to unseen series. The agent explores development series using code and submits candidate detector programs for evaluation. The framework evaluates each candidate on validation data, returning aggregate metrics and selected diagnostic cases through which the agent can inspect validation signals and labels. This bounded feedback is designed to help the agent diagnose failures, explore alternatives, and refine its code while reducing the risk of overfitting to the validation data. The framework retains the best-scoring detector program, while holdout series are reserved for final evaluation. We further train the agent with reinforcement learning in the same development environment, improving its ability to explore data, construct detector programs, and use diagnostic feedback to guide subsequent code revisions. Across multiple datasets and evaluation settings, detectors developed by AgentAD outperform established time-series anomaly detection methods. Under matched development budgets, reinforcement learning enables the trained agent to match or even surpass larger general-purpose language models at detector development. The agent framework, training pipeline, and model checkpoints are available at https://anonymous.4open.science/status/AgentAD_anonymous-9D20.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.