acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking LLM abilities on postdoc-level natural science problems

Abstract

Scientific problems have no answer key, so evaluating an agent means deciding what counts as a better solution before any agent is run. Existing benchmarks either match a symbolic answer against a key or, where they require executable work, score it by unit tests, by a rubric, or by position on a competition leaderboard. None of these is the criterion by which a field decides that one published method beats another. STEMBench adopts that criterion directly. Each of its twelve research problems in high-energy physics, astrophysics and neuroscience is scored by its own field's figure of merit, and each is bounded by two measured anchors: a submission that uses no domain knowledge, and an expert reference reproducing a published method, both obtained by running code shipped with the benchmark. The interval between them states how much of a problem is crossed by scientific insight rather than by engineering. Across eight frontier models, task-level wins and domain-level leadership dissociate: the model with the most task wins leads no domain mean. The expert anchor is reached on eight of twelve tasks, while the tasks demanding calibrated uncertainty or causal latent structure resist. A large share of zero scores is structural: the episode ended without a valid submission at all, a failure mode that answer-matching benchmarks cannot see. We release the tasks, graders, anchors and harness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.