Self-Regression-Test:Continuous Integration for Self-Evolving Agents That Rewrite Their Own Skills
Abstract
**ABSTRACT** Self-evolving agents adapt to new tasks by rewriting reusable skills, and every such rewrite is a potential regression: a change that improves one use can silently break behaviors that other tasks still depend on. Aggregate safeguards—parameter regularization, replay, or diff-magnitude penalties—leave the specific behaviors each candidate edit puts at risk unspecified. We study capability preservation at the level of the skill artifact. Each skill carries a versioned regression suite whose predicates are deterministically executable on cached inputs and act as a test-backed, coverage-conditioned partial contract for future edits. Self-Regression-Test (SRT) realizes this contract in three steps. First, a synthesizer converts a skill's past successful and failed trajectories into a fixed suite of typed predicates: outcome, trace, and negative, each carrying a severity label. Semantic outcome checks use a five-vote LLM-judge ensemble whose fallibility is measured on an independent audit rather than assumed away. Second, an executor evaluates every candidate update against the current suite and returns predicate-level verdicts. Third, a controller turns those verdicts into one of three inspectable actions—commit, bounded repair, or rollback. SRT never touches the policy weights or the upstream evolution operator, and refreshes the suite periodically as usage drifts. Across four evolution operators and three lifelong benchmarks, SRT raises retained accuracy by \(8.3\)–\(11.4\) points at a \(0.8\)–\(1.5\)-point cost in adaptation accuracy, at approximately verifier overhead. A six-cell cross-matrix attribution against replay, one-shot Self-Verify, and historical failure-set rerun shows the gain is not explained by rejection alone or extra computation, and an independent -update mutation audit reaches F1 \(0.90\) with hidden-paraphrase pass ( for snapshot exact-match). The guarantee is scoped: it certifies preserved behavior only on inputs whose predicates deterministically cover, LLM-judge errors are audited empirically, and per-skill contracts do not compose across simultaneous multi-skill refactorings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.