Beyond Isolated Tasks: Benchmarking and Automating Full-Cycle Time Series Analysis
Abstract
Real-world time series analysis is a complex, multi-stage process that typically requires human experts to explore heterogeneous records, model temporal patterns, and support downstream decisions. Yet, existing time series models and benchmarks largely focus on carefully curated, isolated tasks, leaving dependencies between stages largely unexplored. We argue that reliable downstream performance requires both accurate upstream results and the ability to effectively leverage them. To address the scarcity of data for evaluating the cross-stage dependencies in full-cycle time series analysis (FC-TSA), we collect real hydrologic records across 15 gauged catchments and introduce the first benchmark for evaluating Full-Cycle Analysis Capabilities for Time Series: FacTS-Bench. It follows the complete real-world hydrologic analysis process, containing over 1.5K scored instances across four core capabilities: (1) Environment Exploration prepares analysis-ready series from unfamiliar workspaces for subsequent stages; (2) Situation Perception identifies events, assesses anomalies, and establishes current conditions; (3) Forecast Orchestration builds forecasts to inform; (4) Warning Decisions produces evidence-backed warning and dispatch plans. Crucially, FacTS-Bench traces how intermediate results affect downstream outputs, capturing cross-stage dependencies overlooked by existing task-specific benchmarks. To further understand cross-stage dependencies, we present AutoTSer, an autonomous analyst for FC-TSA. It coordinates a Planner, Worker, and Auditor through two mechanisms: (1) Evidence-Grounded Collaboration checks intermediate results against source data, while (2) Boundary-Constrained Progression checks whether those results meet the requirements of the next analytical stage. AutoTSer further learns from alternative operations and their downstream effects, using offline reinforcement learning to train a compact controller. This enables (3) Lightweight Self-Evolution without fine-tuning the underlying language models. Full-cycle evaluations demonstrate improved downstream decision quality through more reliable use of intermediate results. Further analysis shows that agent self-evolution helps retain beneficial operations and avoid harmful ones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.