The Evolution Strategies Paradox: Better Models from Worse Perturbations
Abstract
How is it possible for perturbation-based methods like Evolution Strategies (ES) to optimize billion-parameter large language models with population sizes of thirty? Classical evolutionary methods are highly dependent on finding good candidates within generated populations. However, when performing ES in various fine tuning tasks for large language models, a phenomenon can occur where all population members consistently perform worse than the current model, yet with each iteration the model consistently improves. This paradox occurs across models, tasks, and even domains. This paper first reports this ES paradox, then examines how the local reward landscape as well as the ES estimator properties contribute to this paradox: in high-dimensional spaces, exploration leads to worse performance, but still highlights good directions; aggregating those directions then allows making local improvements. These findings provide guidance for taking advantage of perturbation-based methods in LLM fine-tuning and in high-dimensional optimization in general.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.