What to Try, Where to Try It: Online Adaptive Search Policies for Machine Learning Agents.
Abstract
When should a machine learning agent start over, repair a failed program, or refine an existing solution? And which of its own past solutions should it work on? These two decisions determine how an agent spends its experimental budget, yet common search controllers fix both using hand designed rules that do not learn from the outcomes of previous expansions. We propose Factorized Adaptive Search (FAS), an online controller that separates operator allocation, meaning what to try, from parent selection, meaning where to try it. A fitness rate rank bandit chooses among Draft, Debug, and Improve, while operator specific LinUCB models select eligible parents using validation and search history features. Both components learn from a shared validation improvement reward normalized by action cost, start afresh for every run, and never observe held out test feedback. We evaluate FAS as a controlled intervention on top of AIRA DOJO across 11 MLE bench Lite tasks. We hold the language model, prompts, and operators fixed while varying the two policies independently and together. Learning the operator policy alone (FAS OP) raises anytime percentile rank AUC from 0.781 to 0.820 over the fixed AIRA baseline, while using 26% fewer tokens, producing more search valid candidates (42.7% to 52.8%), more medal worthy solutions among them (36.8% to 48.8%), and achieving the best mean final rank across the four conditions (1.59 versus 2.14). Adding online parent selection, either alone or together with the operator policy, instead reduces performance below the fixed rule baseline. This asymmetry is informative: under the sparse and noisy validation rewards available online, learning what to try is the more valuable and more reliably learnable decision, while learning where to try it requires more care than a direct application of contextual bandits provides.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.