Can Self-Evolving Agents Recover from Their Own Mistakes?
Abstract
Self-evolving agents modify their own prompts, workflows and parameters and select modifications by a score. Score-based selection guards against modifications that lower the score, but not against those that remove something the agent needs while the score stays high. We ask whether such loops can recover from these silent mistakes. Without its archive-based parent selection, a stock workflow optimizer deletes its agent's input placeholders, turns every exception into a graceful exit, and replaces its model call with a nonexistent method; the agent never calls a language model, yet its validation score equals the initial agent's, because a third of validation tasks pass when the agent does nothing. Two other optimizers fail differently; only the archive protects the deployed version. In a controlled IFEval-style task with fifteen verifiable conventions, reporting how many conventions an output satisfies prevents the losses that pass/fail feedback causes; but once a convention is lost, neither counts nor anonymous identifiers of the failing convention restore it in any of four rewriters, while supplying its statement does. We prove that feedback which only evaluates the current state cannot repair a loss faster than the rewriter can guess the lost content. An unseen tool outage is diagnosed only when error traces reach the optimizer, and the self-written exception handler turns the fault's score drop into a score gain. In self-training, a small model collapses to a constant answer that an accuracy gate cannot see. We argue for evaluating self-evolving systems by whether their mistakes remain detectable and recoverable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.