Learning the ARTS of Search for Auto ML
Abstract
Automated machine learning (ML) research can be formulated as an iterative search over hypotheses and their code implementations. Contemporary methods navigate this space with heuristics such as MCTS that rank nodes by score alone, conflating the merit of a hypothesis with the quality of its execution. As a result, a promising hypothesis with a preliminary implementation may rank below a modest one that is well tuned. We propose Agentic Reasoning for Tree Search (ARTS), where a reasoning language model inspects prior execution logs, diagnoses whether failures arise from faulty implementations or poor hypotheses, and selects which hypothesis to build on next. As the search progresses, the model suffers reasoning degradation due to longer contexts. In response, ARTS* learns from the previously explored hypotheses using test-time RL. Across 22 tasks from MLGym and MLEBench, ARTS outperforms leading methods by a mean normalized score of 12.0%. With test-time training, a Qwen3-4B scientist matches or exceeds an o3 scientist on 6 of 12 MLGym tasks. Notably, on MetaMaze, ARTS* learns to rediscover the human-best recurrent-memory solution that heuristic methods prune away.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.