acceptodds
Under review as a conference paper at ICLR 2027

EvoForge: Agentic Online Hyperparameter Tuning for Pretraining

Abstract

Pretraining hyperparameters are typically fixed before training begins, while the optimal configuration drifts as training proceeds. Existing automated methods, such as random search, Bayesian optimization, multi-fidelity scheduling, and population-based training, presuppose multiple trials or concurrent replicas; a single large-scale pretraining run yields one trajectory, precluding a second trial. We present EvoForge, to the best of our knowledge, the first system that places LLM agents in the control loop of a running distributed pretraining job. We cast online tuning as a single-trajectory, partially observable sequential decision problem, where observations are induced by past actions, and organize the system around three cooperating agents, namely, one for tuning, one for safety, and one for cross-task reflection. EvoForge distills rules into a two-layer experience store only after sufficient evidence has accumulated. In GPT-2 124M trained on 10B FineWeb-Edu tokens, EvoForge reduces the final validation loss by 4.15% relative to a tuned cosine schedule.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.