Continual Reinforcement Learning with Neuroevolution
Abstract
Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms, such as evolution strategies (ES) and genetic algorithms (GAs), that search directly in weight space through mutation and selection over a population of neural networks. We compare ES and a GA against state-of-the-art continual RL variants and population-based RL, with no method informed of a task switch, across different kinds of environments and environmental changes and policies ranging from a few hundred parameters to million-parameter networks. We find that ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method. To explain this, we study the return landscape around each method's solutions. We observe that ES finds the widest neighborhoods, regions of weight space in which perturbed policies still solve the task, and that the size of the overlap between the neighborhoods of consecutive tasks correlates with the stability-plasticity trade-off of a method. The GA has narrower neighborhoods and moves further per task, so it is more plastic but forgets more. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE suggesting that they are specific to gradient-based optimization rather than non-stationarity. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.