acceptodds
Under review as a conference paper at ICLR 2027

Learning to Evolve: Skill–Policy Co-Evolution for Self-Evolving Agents

Abstract

Agent self-evolution has garnered increasing attention, where LLM agents improve their policies by extracting and internalizing reusable skills from their own interactions. However, we find that current methods can exhibit evolutionary stagnation, where policy improvement gradually diminishes over successive evolution rounds as skill-extraction quality degrades and skill extraction becomes increasingly misaligned with policy optimization. In this paper, we introduce a skill-policy co-evolution method (SCALE) that evaluates candidate skills by their effects on long-term policy evolution. Specifically, SCALE learns a skill-conditioned velocity field to model how each skill transforms the trajectory distribution and constructs an evolution-aware reward based on the Wasserstein distance to a reward-reweighted target distribution. It enables the agent to continuously extract skills aligned with its evolving policy and sustain improvement across evolution rounds. We evaluate SCALE on nine challenging benchmarks spanning scientific reasoning, coding, agentic tasks, GUI control, and real-world environment learning. The results show that SCALE consistently surpasses existing methods across diverse model scales and exhibits strong active learning capability from interaction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.