acceptodds
Under review as a conference paper at ICLR 2027

Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents

Abstract

Experience-driven self-evolving agents aim to overcome the static nature of large language models by distilling reusable experience from past interactions, thus enabling adaptation to novel tasks at deployment time. This process places substantial demands on the foundation model's capacities for abstraction, generalization, and in-context learning. However, existing methods either rely on external system-level designs or optimize only isolated components, without jointly improving the underlying model's capabilities to extract and utilize experience. To this end, we propose **Evolving-RL**, an algorithmic framework that jointly improves experience extraction and utilization. Specifically, we center the learning process on experience extraction and evaluation, using the two supervisory signals derived from evaluation to optimize the extractor and solver separately and thus enable their coordinated co-evolution. Experiments on ALFWorld, ScienceWorld, and Mind2Web demonstrate that Evolving-RL substantially improves experience extraction and reuse, enabling more effective experience-assisted out-of-distribution (OOD) adaptation. Under the same skill-conditioned protocol, it achieves 98.7% relative improvement over GRPO on ALFWorld unseen tasks and 35.8% in overall action accuracy on Mind2Web, together with a 6.3-point OOD score improvement on ScienceWorld. Furthermore, Evolving-RL also serves as an experience-augmented RL algorithm: the Evolving-RL-trained Qwen2.5-7B-Instruct model substantially outperforms its GRPO counterpart on ALFWorld and Mind2Web even without test-time skill injection, demonstrating benefits for both explicit skill reuse and the underlying policy. The code will be released after the review process.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.