Zero-L: Tri-Role Co-Evolution Framework for Self-Evolving LLMs
Abstract
Self-evolving large language models (LLMs) can improve through reinforcement learning on self-generated tasks and feedback, reducing reliance on human-annotated training data. However, existing self-evolution frameworks face a bottleneck of training gain stagnation after first few iterations. This is because of the reward accuracy declination in later iterations. We introduce Zero-L, a framework for long-horizon self-evolution of LLMs. Starting from a single base LLM, Zero-L initializes three independent roles - a Questioner, a Solver, and a Thinker. Questioner and Solver are optimized through reinforcement learning. Questioner is rewarded for providing tasks which challenge the edge of the solver's ability. After questioner update, solver answers the tasks that questioner generates and gets Solver Reward for solving tasks, but this reward is examined to be insustainable. And that's why Thinker-Solver Reward comes onboard, reasoning process of solver's answers is fed into thinker and better process gets a higher thinker-solver reward. At the same time, with high solver reward accuracy in early iters, thinker is trained to rate reasoning process that leads to right answers. But this indicates that the ability of thinker is weak in early iters, and that's why we introduce Iteration Dependent Reward Policy (IDR Policy) which mix solver reward and thinker reward through one math function which gives solver reward a higher weight in early iterations and a lower weight after accuracy declination. Empirical results show Zero-L outperforms strong baseline at +5.7 % in math reasoning benchmarks and +6.4% in general reasoning benchmarks. More importantly, the mean IDR reward for correct responses remains above 60% across all ten iterations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.