EvolAudit: An Evaluation Framework For Measuring, Interpreting, and Improving Self-Evolving Agents
Abstract
Self-evolving agents autonomously improve their capabilities by updating their components based on available feedback. They are typically evaluated by the final task performance. However, aggregate performance may hide performance drops on individual instances even when the overall score improves. Analyzing these performance declines provides fine-grained guidance for improvement. To fill this gap, we propose an audit framework EvolAudit that measures evolution, interprets performance changes, and provides guidance for improving self-evolving agents. During evolution, EvolAudit tracks instance-level changes between iterations, recording a gain when an instance fails in one iteration but succeeds in the next and a regression when the opposite occurs. To assess evolution quality, we define Evolution Efficiency as the proportion of gains among all gains and regressions. For evolution interpretation, we examine component updates at each iteration, and then relate them to instance-level gains/regressions to assess their contribution to performance changes. Building on this analysis, EvolAudit translates instance-level diagnostic insights into targeted component improvements to increase gains relative to regressions. We apply EvolAudit to harness evolution on terminal tasks and memory evolution on web tasks. Experimental results reveal frequent instance-level regressions throughout evolution and show that some updated components are rarely activated during execution. Guided by these findings, targeting previously untouched components improves final performance from 61.77% to 72.06%, showing that EvolAudit can effectively identify evolution failures and guide further improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.