DiagMem: Experience-Grounded Diagnostic Memory for Agentic Time-Series Forecasting
Abstract
Agentic time series forecasting extends standalone forecasting by incorporating tool use, feedback, and historical experience into the forecasting process. However, reusing this experience to diagnose a new forecast requires relating verified error feedback from completed forecasts to a current forecast whose outcome is still unknown. We introduce DiagMem, an experience-grounded diagnostic-memory framework for agentic time series forecasting. Its key idea is to interpret disagreement among heterogeneous forecasters using verified error patterns from similar diagnostic states. DiagMem organizes completed forecasting cases into diagnostic prototypes and uses a large language model to summarize their aggregated forecast, disagreement, and error evidence as structured scores and semantic descriptions. For a new window, the historical input, base forecast, and cross-model disagreement define its diagnostic state, which guides prototype retrieval. A lightweight numerical executor combines the retrieved diagnostic knowledge with current numerical inputs to estimate trend and seasonal residual corrections while keeping the base forecaster frozen. Across seven multivariate benchmarks and four forecast horizons, DiagMem delivers broad improvements over diverse frozen base forecasters, while ablations support the contributions of cross-model disagreement, state-conditioned diagnostic memory, and the joint use of structured scores and semantic descriptions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.