IdeaCompass: Making Exploration Measurable in Autonomous AI Research Agents
Abstract
Autonomous machine learning (ML) research agents conduct research by proposing ideas, implementing them, and running experiments. Yet existing autoresearch systems lack an explicit quantitative account of the idea landscape they have explored: they do not directly measure how far apart past ideas are, in which directions the search has moved, or how broadly the search has explored the idea space. General-purpose text embeddings do not fill this gap. We find that they do not reliably capture the distance and relative direction between ML research ideas, placing a change in training duration closer to a learning-rate change than to another duration change. We introduce IdeaCompass, a quantitative map of ML ideas: an ML-aware embedding whose vector direction encodes what an idea changes and whose norm encodes how substantial the change is. On this map, an agent's research history becomes a measurable trajectory, from which IdeaCompass diagnoses whether stalled search is over-concentrated (collapse) or over-dispersed, and steers subsequent proposals toward diversification or consolidation accordingly. Across eight FML-bench-Lite tasks, IdeaCompass improves both Autoresearch and AdaptiveSearch, which become the top two of nine evaluated agents; Autoresearch's pairwise win rate rises from 0.52 to 0.71. The representation achieves 0.740 AUROC on direction discrimination and 0.72 Spearman correlation between norm and modification scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.