Why Do Self-Evolving LLM Agents Fail? A Study of Memory-Based Learning at Test Time
Abstract
Memory-based self-evolving LLM agents promise continuous improvement by reusing experience without updating model weights. Yet a failure can originate in earlier learning, memory retrieval, or task execution, and outcomes alone cannot tell which. We introduce WTF (What's the Failure), a human-calibrated framework that diagnoses failures by linking task trajectories to memory artifacts and update histories, and tests selected diagnoses by replaying tasks from saved memory state. Across 9,544 traces spanning four self-evolution methods, five benchmarks, and three base LLMs, we derive a taxonomy of 25 failure classes. Most failures reflect persistent model, environment, or harness limitations, but 10.9% are memory/state-induced: a *memory tax*. This tax is heaviest in methods that inject their entire memory into every task, which show nearly three times the memory-induced share of methods that retrieve a few entries, and the gap widens as memory grows. Lessons from successful attempts can discourage useful actions, and correct knowledge can be missed at retrieval or applied beyond its scope. In controlled reruns, removing harmful guidance restores useful behavior without adding knowledge; across five selected cases, a memory-free agent raises pooled success from 1/18 to 13/18. Yet fixing a targeted behavior does not guarantee task success, and a fix for one task may not transfer. Reliable self-evolution thus requires effective learning, appropriate retrieval, grounded execution, and trustworthy evaluation. We will release code, memory artifacts, and task traces.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.