EvE: An Alternate Optimizer to Adam
Abstract
Adam and its variants dominate neural network training, but a single training run only reveals whether one configuration works well after most of its budget is already spent. The concept is a poor fit for hyperparameter or architecture search, where many configurations must be ranked cheaply and pruned early. As an alternative to Adam, we introduce EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one candidate via the DE operator. EvE only runs a short burst of gradient descent on it if the DE step fails to improve on the incumbent. Selection is greedy, so on a deterministic objective the best-so-far value is provably monotone non-increasing, and because gradients are used only as a targeted rescue, per-iteration cost stays within a constant factor of a single Adam step regardless of problem dimension. Under a fixed, evaluation-cost-matched budget, EvE wins or ties Adam on 76% of 70 (problem, dimension) cells across seven classical scalable benchmarks scaling up to one million variables. On three real neural-network training tasks (an MLP on MNIST, and LoRA fine-tuning of a 1.5B-parameter language model on two datasets) EvE consistently finishes the same charged budget 1.7 to 3.9x faster in wall-clock time, at a modest cost in final quality (about one accuracy point on MNIST, roughly 9-11% higher relative test loss on the two language fine-tuning tasks; on GSM8K's full test set, Adam is also about 5 accuracy points more accurate at matched budgets, and fine-tuning lowers accuracy below the base model for both optimizers). Inside successive halving on UCI Adult Income, EvE completes hyperparameter and architecture searches 3.1-3.5x faster, and its ranking of candidate configurations agrees with Adam's about as well as Adam's agrees with itself across training seeds (Kendall's tau of 0.66-0.69). These results position EvE not as a total replacement for Adam as a final-stage trainer, but as a fast, gradient-aware proxy for the search-heavy, budget-constrained regime one level up: ranking candidate configurations, or pruning clearly bad ones early, where a fast, approximate signal is worth more than a slow, precise one. Code will be made publicly available, with a link, in the final version.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.