AIS-Bench: An AI Scientist Benchmark for Autonomous Scientific Discovery
Abstract
The emergence of AI Scientists—agents that generate hypotheses, implement methods, and report findings—calls for evaluation beyond isolated tasks or final manuscripts. Existing benchmarks emphasize individual capabilities or end-to-end outcomes, leaving the fidelity of intermediate scientific artifacts underexamined. We introduce AIS-Bench, a diagnostic framework that evaluates automated research through three independently scored phases: Reference-Conditioned Idea Reconstruction, Code Implementation, and Controlled Reporting. From 3,174 machine-learning papers, our construction protocol yields 753 aligned <Context, Idea, Code, Paper> quadruplets with executable reference implementations. We evaluate 13 prompting strategies, specialized systems, and end-to-end agents using phase-specific metrics, with judge-based components validated against held-out human judgments. On matched quadruplets under a common human-calibrated artifact-validity criterion, implementation receives the lowest phase score; paired bootstrap intervals and alternative aggregations preserve this ordering. Audits further show that functional-logic deviations dominate classified implementation failures, while metric exaggeration and unresolved citations persist in reports despite specified evidence. AIS-Bench enables systematic, traceable diagnosis of scientific-fidelity degradation across the research lifecycle and supports the development of verification-centered AI Scientists.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.