EvoForge: Agentic Online Hyperparameter Tuning for Pretraining
Abstract
Pretraining hyperparameters are typically fixed before training begins, while the optimal configuration drifts as training proceeds. Existing automated methods, such as random search, Bayesian optimization, multi-fidelity scheduling, and population-based training, presuppose multiple trials or concurrent replicas; a single large-scale pretraining run yields one trajectory, precluding a second trial. We present EvoForge, to the best of our knowledge, the first system that places LLM agents in the control loop of a running distributed pretraining job. We cast online tuning as a single-trajectory, partially observable sequential decision problem, where observations are induced by past actions, and organize the system around three cooperating agents, namely, one for tuning, one for safety, and one for cross-task reflection. EvoForge distills rules into a two-layer experience store only after sufficient evidence has accumulated. In GPT-2 124M trained on 10B FineWeb-Edu tokens, EvoForge reduces the final validation loss by 4.15% relative to a tuned cosine schedule.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.