Evidence-Driven Discovery Loops for AI-Train-AI: From Trials to Better Decisions
Abstract
An emerging frontier of automated AI research is AI-train-AI, in which agents autonomously propose, implement, and evaluate machine learning solutions within a closed discovery loop. These experiments consume substantial resources and often fail or yield marginal gains. Under a limited budget, final solution quality depends on choosing when to refine or combine implementations and which candidates to build on. Each trial produces evidence beyond a scalar score, yet this evidence is often only partially used when operations and sources are selected. We introduce the Evidence-Driven Discovery Loop (EDL), which treats evidence use as part of algorithm development. It connects accumulated trial evidence to subsequent search decisions. Its outer loop captures successful and failed trials, deterministically organizes implementations by algorithmic direction, and links modification intents to evaluated outcomes. Its inner agentic policy, Operation-aware Evidence Control (OEC), starts from compact direction cards and operation-level meta signals, actively queries supporting histories and code, and constructs plans specifying operations, source roles, changes, and execution allocations. Both levels serve final solution quality within a budget that accounts for decision and execution costs. Successful and failed attempts update the evidence available to later decisions, making trial and error cumulative without changing controller parameters. Experiments on machine learning engineering tasks show strong solution quality across three model backbones, and NatureBench Lite results extend the comparison to scientific algorithm development. Selected component interventions characterize organization, action control, and evidence acquisition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.