EffectPreconception: Guiding ML Research Agents through Effect Prediction and Observation
Abstract
Autonomous machine-learning research agents optimize models through iterative code modifications, relying on scalar validation scores to guide exploration. A validation score reveals whether performance changed, but not whether the edit achieved its intended effect, leaving agents unable to separate flawed ideas from uncalibrated implementations. We introduce EffectPreconception, a framework that separates the predicted effect of an edit from its performance outcome. Before execution, the agent states an explicit, testable prediction about the observable behavior the edit should induce. When domain decompositions or accumulated observations support it, an outcome relation maps this predicted effect to an expected change in validation score. After execution, EffectPreconception checks the measured effect against the prediction and records the result in a persistent evidence ledger that informs every subsequent proposal, including evidence from discarded candidates. We evaluate EffectPreconception on Autoresearch, MLE-bench Lite, and PostTrainBench. On Autoresearch, it reduces mean best validation bits per byte from 0.9501 to 0.9404 across three runs with Sonnet 5. Applied to MLEvolve, it increases the mean gold-medal rate on MLE-bench Lite from 56.1% to 65.2% with Opus 5. These results demonstrate that evaluating targeted structural effects alongside validation performance provides substantially better guidance for autonomous research agents than score-only feedback or direct score prediction. Code: https://anonymous.4open.science/r/EffectPreconception-8022.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.