-Evolve: LLMs Evolve Their Own Generator Programs for Self-Improvement
Abstract
In group-relative reinforcement learning with binary verifiable rewards, all-correct or all-incorrect response groups contribute no reward-driven policy gradient. Sustained self-improvement therefore requires problem instances that continue to yield both correct and incorrect responses as the Solver changes. We introduce , which retains generator programs in an archive, each representing a task family of problem instances sharing the same structure. Each iteration, the Solver reassesses retained programs on newly sampled instances and, as the Evolver, mutates them into child programs. Reasoning Quality () estimates the current usefulness of retained programs and their children by multiplying learnability, measured as verifier reward variance, by uncertainty, measured as mean token entropy. We motivate this product by deriving an upper bound on the squared norm of a centered-reward surrogate gradient in output-logit space. Fitness ranks eligible retained programs for training, while retention confines competition to cells defined by mathematical domain and output form, leaving programs in other cells available for later reassessment and mutation. With Qwen3-4B-Base and Qwen3-8B-Base, achieves final mathematical benchmark gains of and percentage points over the base models. Both scales reach their highest averages late in training, and the final 8B checkpoint exceeds R-Zero's peak by points. Classified generated instances remain broadly distributed across domains and output forms late in the 8B run.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.