acceptodds
Under review as a conference paper at ICLR 2027

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Abstract

Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). This paper systematically investigates the behavior and design of ES post-training. we : broader reasoning coverage that better exploits the reasoning capabilities of pretrained LLMs. Empirically, ES generally improves Pass@1 over the base model and achieves higher Pass@ than GRPO, which can exhibit entropy collapse. Theoretically, we derive sufficient conditions under which an ES update improves the center policy's Pass@. We further propose a sequential ES–GRPO training method in both orders, which can combine GRPO's Pass@1 strength with ES's Pass@ gains. we find that despite substantial parameter drift, removing most small coordinate updates of ES largely preserves target-task Pass@1 performance. Held-out evaluations further show that this large drift need not imply broad forgetting. we study how ES design choices affect its effectiveness, finding that comparable training rewards can be maintained with smaller ES populations in larger LLMs. These findings position ES as a distinct post-training paradigm for LLM reasoning rather than a less effective, memory-efficient alternative to GRPO. Our code is available at https://anonymous.4open.science/r/understanding-es-3D73.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.