acceptodds
Under review as a conference paper at ICLR 2027

Building the Reliable Evaluation Layer of Auto Research with On-Demand Benchmarks

Abstract

Auto Research is emerging as a new paradigm that extends Large Language Model (LLM)-Agent assistance to the autonomous exploration and discovery of scientific experiments. However, research on novel topics may lack reliable benchmarks for supporting its claims. While Auto Research systems generate test cases to fill this gap, the resulting evaluations may favor their own methods, raising concerns about fair comparison. To this end, we introduce Idea2AutoEval, a general framework that constructs reliable evaluations on demand from research intents for Auto Research. The framework iteratively organizes prior benchmarks into domain taxonomies to identify established evaluations and clarify the scope of new ones. Building on the in-domain benchmark library, it combines research claims and benchmark coverage to define the evaluation scope and scoring criteria. To support evaluation integrity, we adopt the information separation mechanism that exposes only task inputs to evaluated models while reserving reference data exclusively for the evaluator. Guided by these, our framework couples benchmark generation with iterative quality assurance to support the validity of model comparisons through semantic verification and alignment with research goals. Experiments span eight benchmark settings across four task categories, with datasets produced by two generation models. Spearman correlations with public model rankings exceed 0.8 in 14 of 16 comparisons, while rankings from the two generators correlate above 0.85 in seven of eight settings. On four representative tasks, Idea2AutoEval outperforms BenchMaker and BenchBench in macro-average Spearman correlation with public rankings, reaching 0.8810. Further analyses show that feedback can improve agreement with public rankings, while calibration of the intended difficulty remains task dependent. These findings support generated benchmarks for model comparison and suggest their potential as evaluation feedback for Auto Research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.