AlphaIdeaBench: Benchmarking Automated Research Idea Generation under Feasibility Constraints with Peer Review Alignment
Abstract
Research ideas are always proposed under resource constraints: a plan must fit the available compute, API credit, and time. Yet evaluations of research-ideation agents mostly score unconstrained novelty or overlap with a published paper, and their LLM judges are rarely checked against peer review. We introduce AlphaIdeaBench, a benchmark of 100 ideation tasks (50 workshop calls for papers and 50 open problems stated in surveys), each paired with an explicit resource budget and a shared proposal schema. An LLM judge scores each proposal on 12 dimensions covering validity, novelty, and feasibility. To validate the judge, we compare its scores on four novelty axes (problem, method, scenario, and insight) with scores derived from reviewer comments on 600 papers from ICLR 2026, NeurIPS 2025, and ICML 2025. For problem and method novelty, Gemini 3.7 Flash is within one point of the reviewer-derived score on 94–97% of papers; agreement on scenario novelty is relatively lower. We evaluate eight published harnesses on a shared backbone and, within one harness, eight backbone LLMs. CoI-Agent obtains the highest overall score, narrowly ahead of SciMON, which leads on method and insight novelty. All harnesses score low on scenario novelty and rarely change the problem setting. The backbone largely determines which open problem a proposal targets, whereas the harness determines whether a proposal addresses one problem or combines several. Weaker backbones such as Llama 4 Scout fail to carry out that harness's strategy and score low on budget feasibility. Finally, when only the budget is varied on three topics, a larger budget changes the experimental setting on compute-intensive topics but only adds tests on an inexpensive one. We release the tasks, rubrics, and scores to foster standardized, reproducible benchmarking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.