Dimension Is Free, Horizon Is Not: When Evolution Strategies Rival Policy Gradients for LLM Fine-Tuning
Abstract
Policy gradient (PG) methods such as GRPO are the dominant approach to LLM fine-tuning, but they suffer from imprecise credit assignment, a limitation that becomes especially pronounced in long-horizon agentic tasks. Recent work has established Evolution Strategies (ES) as a scalable alternative to PG. Because ES evaluates each perturbed model by its total return, it avoids per-token credit assignment, which makes it a natural candidate for long-horizon settings. We analyze ES and PG within a single framework, as two estimators of the same gradient that explore in different currencies. ES perturbs parameters and therefore pays a dimension tax, whereas PG samples actions and therefore pays a horizon tax. We show that the taxes actually incurred are effective quantities. For ES, the tax is , the number of reward-relevant parameter directions. For PG, it is , the number of token positions that contribute reward-relevant variance, which increases only when reward-irrelevant tokens share both credit and weights with reward-carrying ones. To disentangle the two, we construct a synthetic task in which and can be controlled independently. On this task, PG outperforms ES at high and short , whereas ES outperforms PG when is long relative to , and this region expands with the training budget of ES. We then estimate proxies for and on 27 NLP, mathematical reasoning, and agentic tasks with Qwen2.5-7B-Instruct and relate the resulting atlas to head-to-head ES–PG comparisons. Our results support ES as a scalable alternative for LLM fine-tuning, particularly for long-horizon tasks with sparse credit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.