acceptodds
Under review as a conference paper at ICLR 2027

Evolving Together, Failing Apart: The Skill-Policy Co-evolution Trap in Self-Improving Agents

Abstract

Skill–policy co-evolution—in which an agent’s policy model and skill bank improve each other during training—is a promising paradigm for self-improving agents. Current evaluations assess the co-evolved policy on the final bank with which it was trained (the delivered bank). After deployment, however, skills may be rewritten, added, or removed more frequently than the policy model is retrained, leaving the policy to operate on a drifted skill bank addressed by neither the training objective nor existing evaluation protocols. Through theoretical and empirical analysis, we establish that the fragility masked by this blind spot is structural rather than incidental. The acceptance rule shared by these methods, which we term cooperative acceptance, ties a policy’s competence to the bank with which it co-evolved, inducing fragility under skill drift regardless of implementation; empirically, sensitivity to such drift grows as co-evolution training proceeds. We therefore introduce the CO-evolution Drift Audit (CODA), a deployment-oriented evaluation protocol that simulates skill drift and measures worst-case success rate, consistency rate under meaning-preserving edits, and resistance rate to corrupted skills. Across four co-evolution methods and five benchmarks (six settings), worst-case success rates fall 5–13 points below the peaks attained by the corresponding delivered pairs, erasing 52–93% of the delivered banks’ gains over the no-skill baseline. CODA thus reframes the target of co-evolution—from peak performance on the co-evolved bank to the ability to harness skills however they evolve—and provides a shared standard for auditing future methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.