Discover and Internalize: A Continual Co-Evolving Training Loop for Self-Improvement Agents
Abstract
Agent self-improvement requires not only learning better behaviors, but also discovering what better behaviors could be. Existing approaches typically couple behavioral discovery with policy optimization, limiting the exploration of alternative behaviors and making it difficult to fully exploit potentially useful behavioral strategies. We reframe agent self-improvement as a behavioral hypothesis search problem and view skills as a behavioral hypothesis space: a skill is a candidate hypothesis to be tested, rather than an already validated capability. Based on this view, we propose Discover-and-Internalize Co-Evolution (DICE), a co-evolutionary framework between behavioral discovery and policy internalization. Behavioral discovery evolves and composes skills to explore new behavioral hypotheses, while policy internalization consolidates promising behaviors into the model weights. The updated policy then provides a stronger foundation for the next round of discovery, forming a mutually reinforcing evolution loop. Our results demonstrate that explicitly separating behavioral discovery from policy internalization, while combining them into a co-evolutionary loop, provides an effective paradigm for continual agent self-improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.