acceptodds
Under review as a conference paper at ICLR 2027

Are Your Agents Ready to Evolve? Assessing Evolution-Readiness in Agents Through Paired Safety and Utility Scoring

Abstract

As AI agents evolve at runtime by accumulating memories, saving reusable skills, revising prompts, and delegating to subagents, they are generally expected to improve with experience. This assumes evolution-readiness, that what an agent learns from one task will leave it no less safe and no less capable on the next. Traditional agent-safety evaluations hold an agent's persistent state fixed and score standalone tasks, so they cannot test this assumption. Therefore, we introduce EvoAudit, a task suite purpose-built for agent evolution that uses directed source-target task pairs and matched cold and evolved conditions to isolate how evolution changes subsequent behavior. Tasks span benign and dangerous settings and are scored both on safety and utility. We run controlled experiments across five models using a self-evolving agent framework covering seven evolution mechanisms. We find that most agents exhibit substantial run-to-run variation even without evolution, yet evolution still produces more change than the control baseline. Evolution does not make agents better overall, and when both tasks are dangerous, it doubles the rate at which they trade off utility for safety. We then inject controlled good and bad memories to test the potential effects of evolution. A single injected memory affects safety far more than anything the agents write autonomously, and the models that appear unaffected are the ones that rarely read their memories. We find that the evolution outcome is not predictable from a particular evolution mechanism or model. Our audit suggests that current agents may not yet be fully evolution-ready, as evolution provides limited gains and does not reliably improve behavior across tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.