The Shoulders of Giants: Benchmarking AI agents on open scientific research problems
Abstract
Can AI agents carry out research on open scientific problems? We introduce Shoulders of Giants, a benchmark that provides an agent with a science paper along with a research direction to pursue and tasks it to write a follow-up paper. We work with scientists at US DOE national laboratories, who each contribute an unpublished follow-up direction from their own work together with the necessary conditions a submission must meet. While assessing the quality of a follow-up paper is challenging, using these conditions to form a rubric makes evaluation more tractable for research problems that have no reference answer. We have seven frontier agents write 15 follow-ups each, spanning quantum computing, plasma physics, computational chemistry, materials science, and scientific machine learning domains. None of the 105 submissions satisfies all the conditions, and 99 fail our integrity check, most often because the paper does not match its code or reports numbers that contradict the recorded evidence. Overall, agents readily produce papers that look like follow-ups, but none meets the standard of the domain scientist who proposed the direction, leaving substantial headroom.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.