SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?
Abstract
Autonomous AI research agents aim to speed up scientific discovery by automating the research pipeline from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamental bottleneck whether LLMs can judge the methodological viability of a research idea investing time and computational resources. Without a reliable “first-gate” filter, autonomous agents risk scaling flawed methodology rather than accelerating meaningful science. We introduce , a curated benchmark of 1,099 real research proposals, with source verified extraction, reviewer-derived proxy labels, and proposal-only expert judgments on a subset of 327 proposals. Across 12 frontier LLMs, our evaluation reveals a pervasive : standard prompting yields a mean false-positive rate. Aggressive prompting reduces false approvals but substaintially increases false rejections. Additional analyses show that optimism persists under controls for public corpus contamination, and surface features, as well as in literature-grounded agent evaluation. These results suggest a limitation in judging scientific soundness, indicating that current LLMs are not yet reliable as standalone gatekeepers for scientific rigor.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.