EvoTrace: Benchmarking Self-Evolving Agents on Procedural Tasks through Experience-to-Behavior Tracing
Abstract
Large language model (LLM) agents deployed in professional domains repeatedly encounter requests governed by the same underlying procedure, where a missed operation, conditional decision, or dependency can invalidate an execution. For self-evolving agents that seek to improve from experience, effective evolution requires extracting reusable procedural knowledge from prior task interactions and correctly applying it to new requests. However, existing benchmarks mainly reveal aggregate performance changes or the quality of retained knowledge, leaving it unclear which procedural requirements are actually learned and where the experience-to-behavior pathway breaks down. We introduce EvoTrace [GitHub](https://anonymous.4open.science/r/EvoTrace-20271), a benchmark that uses Agent-Executable Standard Operating Procedures (AE-SOPs) as a common reference to track each procedural requirement across actual experience, explicit retention, and held-out execution. EvoTrace comprises 115 AE-SOPs across seven professional domains and 10,792 executable instances. Experiments across *eleven* representative self-evolution methods show that effective evolution cannot be inferred from either the form or the coverage of retained knowledge alone: similar forms can accompany substantially different gains, while broader coverage does not necessarily translate into more reliable execution. Instead, different mechanisms exhibit distinct bottlenecks in how procedural requirements progress through the experience-to-behavior pathway, and these vulnerabilities vary across requirement types. Targeting the identified bottlenecks with lightweight interventions improves subsequent execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.