Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
Abstract
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). This paper systematically investigates the behavior and design of ES post-training. we : broader reasoning coverage that better exploits the reasoning capabilities of pretrained LLMs. Empirically, ES generally improves Pass@1 over the base model and achieves higher Pass@ than GRPO, which can exhibit entropy collapse. Theoretically, we derive sufficient conditions under which an ES update improves the center policy's Pass@. We further propose a sequential ES–GRPO training method in both orders, which can combine GRPO's Pass@1 strength with ES's Pass@ gains. we find that despite substantial parameter drift, removing most small coordinate updates of ES largely preserves target-task Pass@1 performance. Held-out evaluations further show that this large drift need not imply broad forgetting. we study how ES design choices affect its effectiveness, finding that comparable training rewards can be maintained with smaller ES populations in larger LLMs. These findings position ES as a distinct post-training paradigm for LLM reasoning rather than a less effective, memory-efficient alternative to GRPO. Our code is available at https://anonymous.4open.science/r/understanding-es-3D73.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.