acceptodds
Under review as a conference paper at ICLR 2027

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Abstract

LLM agents are increasingly capable of improving models, systems, and other technical artifacts through autonomous, long-horizon experimentation. Yet existing benchmarks primarily evaluate these agents by final task scores, which capture outcome quality but provide limited insight into the research process. Evaluating runs in isolation also leaves unclear whether accumulated experience improves subsequent performance. To address these gaps, we introduce a unified evaluation framework that follows the research loop to assess how agents formulate directions (Solution Framing), implement solutions (Execution), and preserve or recover progress through feedback (Feedback Control). We further use controlled comparisons within and across tasks to examine whether experience from this loop improves subsequent research. Evaluating seven frontier models across a range of long-horizon AI R&D tasks, we find that similar final scores can conceal distinct process bottlenecks and divergent responses to experience. Some agents struggle to identify effective directions, while others implement strong solutions but fail to retain or restore them after subsequent experiments fail. Likewise, experience reuse can improve one agent's performance while degrading another's, despite comparable baseline performance without experience. These findings provide a more detailed account of autonomous research capabilities and limitations, informing model training, inference-time strategies, and experience management toward more reliable research and self-improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.