What Actually Evolves in Self-Evolving Reinforcement Learning?
Abstract
Self-evolving reinforcement learning (RL) aims to train a model that learns from self-generated tasks, without additional human annotations. It typically alternately trains two roles: a proposer that generates training tasks and a solver that learns from them, spontaneously improving the performance on downstream tasks. However, it remains unclear what drives these gains, and whether the proposer can keep generating tasks that sustain the solver’s improvement. In this study, we systematically examine current self-evolving RL across six methods and six models on seven reasoning benchmarks, with RL on human-annotated answers (human-annotated RL) as a reference. We track how pass@1 and pass@128 change, identify the source of the gain, and test whether later proposers write more useful tasks than earlier ones. Our analysis reveals three findings. (1) Most solver gains come from questions the base model already answers correctly. (2) Self-evolving RL corrects far fewer errors than human-annotated RL. (3) Proposers do not teach more after further training: the same solver does no better, and sometimes worse, when trained on tasks from a later proposer than on tasks from an earlier one. Our theoretical analysis of a single solver update shows that it can only strengthen answers the solver already gives. Although recursive self-improvement aims to go beyond human supervision, current self-evolving RL rarely outperforms human-annotated RL. Taken together, our findings suggest that current gains primarily reflect sharpening, and that self-evolving RL has so far only partially achieved sustained recursive self-improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.