acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Program Search for AutoML: A Reachability Analysis of LLM-Edited Training Programs

Abstract

Do the code edits an LLM agent makes to a training program reach anything a hyperparameter optimizer could not? End scores cannot tell a new mechanism from a renamed knob. We make the question executable: relative to a frozen program, its configuration space, and an extension declared in advance, we sort every edit into four reachability classes by comparing parent and child forward outputs at fixed configurations. Only edits outside the extension can credit program search, and only against a tuner given the same extension. We audit 467 edits a Claude agent proposed to LoRA adapter code and run controlled searches on Pythia-160M and 410M over ARC-Easy, RACE, and OpenBookQA. Of 332 primary edits, 76% are hyperparameter-level: 56% reconstruct exactly inside the extension and 20% rescale one coordinate. At an equal budget of 492 evaluations, neither our controller nor OpenEvolve or ShinkaEvolve separated on validation from TPE on the unedited program. A best-child acceptance rule overstates progress: in our controller's archives, 38 of 51 admitted programs entered because the best score of a short local tuning run beat the parent's by more than one point, below single-run noise, not because the edit helped at matched configurations. Given a wider surface, the agent produced uncovered programs in every seed, yet none showed an established gain over a tuner given the extension; on OpenBookQA that tuner matched the agent's sealed-test gain. Program search should be credited only for uncovered edits, against a coverage-matched control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.