Think Before You Experiment: Efficient Experiment Selection for AI Research Agents
Abstract
Autonomous AI research agents improve solutions through repeated experimentation, but experiments are expensive while LLM inference is comparatively cheap. This asymmetry raises a simple question: which ideas are worth an experiment? Existing agents typically generate and execute candidate ideas sequentially, devoting little computation to deciding which candidate should be tested next. We introduce IdeaMAP, a lightweight plug-in that spends additional inference on experiment selection. IdeaMAP formulates selection as quality-diversity optimization over a semantic space of research ideas, maintaining a MAP-Elites-style archive of executed experiments and prioritizing candidates predicted either to improve existing elites or to fill promising unexplored niches. Across FML-bench-Lite, IdeaMAP improves multiple host research agents and yields agents that outperform existing systems including The AI Scientist v1 and v2, AIDE, and OpenEvolve. IdeaMAP reaches the host agent’s final validation performance using 31% fewer validation experiments and, at the same validation budget, achieves 6% higher final validation performance. These results show that allocating additional inference to experiment selection can improve the efficiency and performance of autonomous research agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.