Learning to Evolve: Skill–Policy Co-Evolution for Self-Evolving Agents
Abstract
Agent self-evolution has garnered increasing attention, where LLM agents improve their policies by extracting and internalizing reusable skills from their own interactions. However, we find that current methods can exhibit evolutionary stagnation, where policy improvement gradually diminishes over successive evolution rounds as skill-extraction quality degrades and skill extraction becomes increasingly misaligned with policy optimization. In this paper, we introduce a skill-policy co-evolution method (SCALE) that evaluates candidate skills by their effects on long-term policy evolution. Specifically, SCALE learns a skill-conditioned velocity field to model how each skill transforms the trajectory distribution and constructs an evolution-aware reward based on the Wasserstein distance to a reward-reweighted target distribution. It enables the agent to continuously extract skills aligned with its evolving policy and sustain improvement across evolution rounds. We evaluate SCALE on nine challenging benchmarks spanning scientific reasoning, coding, agentic tasks, GUI control, and real-world environment learning. The results show that SCALE consistently surpasses existing methods across diverse model scales and exhibits strong active learning capability from interaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.