SciVerge: Objective Evaluation of Scientific Reasoning in Large Language Models
Abstract
Language models are beginning to contribute to scientific work, but their scientific reasoning is rated by human experts or language models rather than measured objectively. Measuring scientific reasoning often requires experimentation to determine whether the ideas a model considered have potential. Mathematics provides an alternative: when a scientific system is described mathematically, as a set of states and the actions that change them, whether an idea would work can be derived from the underlying mathematics. We present SciVerge, a benchmark that describes scientific systems mathematically to construct questions that objectively measure a language model's divergent thinking, the production of many possible approaches to a question, and convergent thinking, the narrowing of those approaches to the best one, the two components of creative thought identified in the psychology of creativity, which proposing and selecting a scientific approach require. Each question is posed in three phases: the model proposes approaches, commits to a subset of them, and attempts the committed approaches until one produces a proof or all have failed. SciVerge mathematically determines whether the approaches that are proposed and selected by a model can succeed, and decomposes divergent and convergent thinking into six metrics to evaluate the model's performance. SciVerge contains 500 evaluation instances drawn from over 100,000 generated across 21 systems in quantum computing, molecular biology, robotics, chemistry, crystallography, and number theory. We additionally construct a generation and screening pipeline that verifies the quality of each instance by proof before admission into the benchmark. Across seven models, the strongest solves 55 percent of instances. Scoring each phase reveals three failure modes, one per phase, that the solve rate alone would not, each calling for a different adjustment. Enabling reasoning improves selection and execution far more than proposal, and every model is far better at disproving a claim than at proving one: the strongest disproves 86 percent of the false claims but proves only 24 percent of the true ones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.