acceptodds
Under review as a conference paper at ICLR 2027

SciGame: A Controlled Benchmark for Evidence-Driven Scientific Law Discovery in Language Model Agents

Abstract

Large language model agents are increasingly applied to scientific discovery, yet existing benchmarks conflate law inference with programming, assess static fits over unseen numerical prediction, and fail to separate empirical induction from pre-trained priors. We introduce SciGame, a controlled benchmark that distills literature into structured scientific cards to instantiate paired Original and dimensionally consistent Counterfactual environments under an identical task interface. Spanning canonical Historical laws alongside a live Temporal collection dynamically curated from post-cutoff peer-reviewed papers, SciGame systematically benchmarks agents across both candidate-guided Identification and open-ended Discovery settings under prior-only reasoning, fixed observation batches, and adaptive sequential experimentation protocols. Evaluating five frontier models reveals three core findings: (1) rooted priors become liabilities once governing relations are altered; (2) adaptive experimentation benefits the strongest models but degrades weaker ones under constrained budgets; and (3) high explanation quality can mask quantitative failures, with exact symbolic formulation remaining the primary bottleneck even when agents reliably interpret empirical evidence. These results demonstrate that while current agents can actively investigate systems, reliably inducing the exact physical laws that govern unfamiliar worlds remains an open challenge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.