acceptodds
Under review as a conference paper at ICLR 2027

What Is Worth Training Next? Evaluation-Guided Evolution for Autonomous Trainers

Abstract

Autonomous training systems can continually adapt training strategies and improve models with limited human intervention. However, their next optimization directions are often chosen based on limited evaluation feedback such as benchmark scores, which can lead to repeated local optimization. More importantly, existing approaches often assess the value of an optimization direction only after completing an expensive training run. They offer little guidance on what is worth training next before the budget is committed. Such guidance is essential when training resources are limited. We introduce EvalEvo, an evaluation-guided framework that turns evaluation evidence into prospective training decisions. Before each costly run, EvalEvo acquires targeted evidence to verify observed progress, diagnose the factors limiting further improvement, and assess which candidate directions can provide useful learning signals. It organizes this evidence in a global directed acyclic graph (DAG), allowing future decisions to draw on accumulated cross-version experience. At each iteration, EvalEvo uses cross-version evidence and the current exploration state to select the next training direction and formulate a falsifiable expectation for its outcome. When the existing evaluation or decision process becomes insufficient, EvalEvo also adapts the training system itself to support continued optimization. Across language-model pretraining optimization and reinforcement learning (RL) for mathematical reasoning and code generation, EvalEvo consistently outperforms existing baselines under fixed experimental budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.