LEAP-Bench: Turning Raw Data into Scientific Claims
Abstract
Scientific discovery depends on turning raw experimental data into defensible claims through quantitative analysis and physical interpretation. Whether AI agents can carry out this process reliably remains unclear. We introduce LEAP-Bench, a benchmark of end-to-end experimental analysis in materials science. We construct 64 tasks from 12 papers spanning multiple instrument modalities, each task targeting a central scientific claim. Agents analyze raw experimental files without seeing the papers' conclusions or procedures. Using reference answers from the source papers, 475 questions assess whether agents extract the required quantities, correctly reason through the analysis, and justify their scientific conclusions. Across three evaluation rounds, Claude Code with Opus 5 achieves the highest task success rate among the six evaluated agents, at (mean standard deviation). Even when agents correctly answer questions about numerical results and intermediate reasoning, their scientific interpretations can remain incomplete. Twelve tasks remain unsolved by all six agents across all three rounds, including eight requiring model fitting or comparison across measurements. These findings highlight persistent difficulties in completing experimental analyses and connecting the results to scientific conclusions. By evaluating the reasoning from experimental data to scientific claims, LEAP-Bench provides a basis for training AI co-scientists with feedback on their numerical estimates, analytical reasoning, and physical interpretations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.