acceptodds
Under review as a conference paper at ICLR 2027

Do LLMs See the Gorilla in the Data? A Benchmark for Off-Objective Scientific Discovery

Abstract

Scientific discovery depends on achieving stated objectives and also on noticing what those objectives leave unstated. We call this ability off-objective discovery, the recognition, while pursuing a given research goal, of an important finding that the same materials support but the goal never mentions. Such discoveries are largely attributed to human serendipity, and existing benchmarks evaluate only whether agents complete the objective they were given. We introduce UncoverBench, a benchmark that measures this ability against verifiable answers. Each of its 193 tasks is built from a recent peer-reviewed paper. A model receives the paper’s tables as the working data of an ongoing study, together with a research objective, and is asked what should be done next. One of the paper’s central findings is never mentioned, yet the tables support it. Every task also carries decoy conclusions that the tables contradict, which catch reports that assert patterns without checking them. Across twelve frontier models, reports complete the stated objective about three times in four but state the unstated finding in only 27.9% of cases. The misses are failures of recognition. In a sample labelled by hand, about two thirds of reports interpret the evidence for the finding, yet fewer than one in four state it, and models that fetch the right tables on their own still miss the finding in about half of the tasks or more. A generic reminder to look beyond the objective raises discovery to 43.5% and leaves decoy endorsement unchanged. Off-objective discovery thus remains a distinct capability that current frontier models largely lack, whether they read all materials at once or fetch the evidence with tools. UncoverBench provides a verifiable foundation for measuring and improving this capability, moving research agents from completing the objectives they are given toward noticing the findings those objectives leave unstated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.