acceptodds
Under review as a conference paper at ICLR 2027

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Abstract

Isaac Newton discovered the law of gravitation through an iterative process of analyzing observed data such as planetary patterns, finding the mechanisms by describing them in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets, and successfully predicted the existence of planet Neptune. As AI agents rapidly advance, whether they can automate such a process has become a major point of contention. To measure this gap, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments, find the mechanism to explain the observations, and eventually be judged by its derivable scientific insights. To measure this gap, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments, find the mechanism to explain the observations, and eventually be judged by its derivable scientific insights. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.