SciProdBench: Evaluating and Improving Scientific Research by AI Agents
Abstract
AI agents can now carry out substantial parts of scientific research, from surveying the literature to running experiments and reporting conclusions. Whether those conclusions rest on synthesized evidence and correct computation, however, remains hard to assess, and automated research can improve only on what it can verify. We introduce \sciprod, a benchmark of 184 research tasks in Biology, Chemistry, Computer Science, and Physics that span the research cycle, from problem formulation to experiment design, analysis, and reporting. Each task pairs a research objective with required deliverables and a task-specific environment. Report rubrics assess evidence synthesis and scientific reasoning, and for tasks that require computation, hidden executable tests check the underlying artifacts. Tasks and their verifiers are derived from the same graph of scientific evidence, so that every requirement and scoring criterion can be traced to its scientific basis. Evaluating nine models and four harnesses, we find that agents meet 90–97% of instruction-following criteria but only 49–57% of evidence-synthesis criteria and that Codex and Claude Code raise scores on reasoning tasks by 10–23 points while lowering scores on execution tasks by 6–9 points. We further introduce \eulerflow, a multi-agent harness that checks reported findings against artifacts and execution records and returns unsupported findings for revision. \eulerflow improves both reasoning and execution tasks with two different underlying models, raising the overall score of Claude-Opus-4.8 from 44.7 to 53.1.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.