EvoLPO: Continual Reinforcement Learning through Policy Evolution
Abstract
Continual reinforcement learning (RL) with verifiable rewards typically resumes each new task from the latest checkpoint. Although recent work reduces forgetting with replay and regularization and ranks starting checkpoints with short adaptation trials, the recovery step that follows adaptation is left out of this choice. This matters because the latest checkpoint may have the highest current score and still be a poor starting point: adaptation changes its capabilities and recovery changes them again. We therefore treat checkpoint selection as a two-stage decision, in which a controller chooses a parent from an archive of past checkpoints, observes the child produced by adaptation and then chooses how to recover earlier capabilities. Both decisions share the utility of the completed child, so each parent is valued by its outcome after recovery rather than before it. Our approach, , keeps the inner RL optimizer fixed and deploys a single policy at every stage. Across mathematics, code-generation and program-optimization streams, EvoLPO improves new-task learning, retention and final score over sequential training and a sequential two-head controller and exceeds tuned replay in final score at equal total compute; the gain over the two-head controller holds across model scales and unseen model families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.