Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training
Abstract
Self-evolving training improves large language models by generating tasks and learning from self-produced supervision. Existing approaches can alternate between training a challenger and training a solver, requiring an additional reinforcement-learning optimization for task generation which is costly in computation and memory. Hence we introduce and study a new self-evolving framework, Direct Self-Evolving Optimization (DEO), which directly samples from optimal challenger distribution and only trains the solver. Under stated regularity and gradient-dominance conditions, we prove that DEO can learn distributionally robust model. The error bound holds over the KL ball induced by the final optimal challenger. Motivated by this conceptual algorithm, we use a finite-step, Metropolis-style mutation walk to simulate sampling task from the optimal challenger distribution and update only the solver. Our implementation uses finite-rollout uncertainty estimates and an approximate acceptance rule. We evaluate the basic method on Qwen3-4B-Base and Qwen3-8B-Base with 2,000 candidate questions per round and a fixed temperature. At iteration three, DEO attains seven-benchmark average accuracies of 47.57% and 53.60%, respectively, compared with 45.67% and 51.77% for the corresponding no-walk ablation and 45.93% and 52.88% for R-Zero. Meanwhile, our method saves memory and compute time compared to R-Zero. These comparisons use the same three-round horizon and nominal solver-update budget; the methods differ in question-pool size and generation cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.