HeuriTrain: Search Organization and Verification in Autonomous Post-Training
Abstract
Autonomous post-training requires an LLM researcher to choose data and training interventions under costly and variable feedback. Evaluating such systems requires understanding what they discover, how they reach their configurations, and whether observed gains recur. We present HeuriTrain, a framework for budgeted experimental search over complete data and training-recipe configurations. Explicit task contracts specify editable components, execution and verification rules, and budgets, while experimental records connect configuration inheritance, referenced results, executed changes, and outcomes. Within this framework, we study global search policies and Factorized Experimental State (FES), which organizes parent selection, visible history, and edits by intervention surface. Experiments on IFEval and Countdown examine checkpoint discovery and researcher decisions, complemented by fixed-configuration retraining and repeated policy comparisons on a deterministic control task. We observe that researchers can draw on results beyond their selected parent, that permitted and executed edit scopes can differ, and that high discovery scores need not persist under retraining. The control task provides evidence of a short-horizon policy effect, with much of the FES–Greedy separation present at the first proposal. Together, these findings motivate evaluating autonomous post-training through its discovered checkpoints, decision process, and repeatability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.