acceptodds
Under review as a conference paper at ICLR 2027

OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences

Abstract

Progress in the foundational theoretical sciences requires resolving questions beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra attains the highest mean judged solve rate of 14.02%. GLM-5.3-Flash also receives solved judgments and a rate of substantive partial progress comparable to several larger systems. Case comparisons connect stronger outcomes to changes in problem representation, general arguments that extend beyond finite evidence, and proofs of the steps needed to complete a solution. By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.