acceptodds
Under review as a conference paper at ICLR 2027

Beyond Selected Breakthroughs: Evaluating Language Models on Randomly Sampled Mathematical Conjectures

Abstract

The last year has seen a shift in AI-assisted mathematics, with frontier language models producing proofs and counterexamples for long-standing open problems. However, curated benchmarks and selected breakthroughs provide an incomplete picture of LLMs performance and the computational resources and problem-specific mathematical guidance needed to achieve reported results. To address this, we introduce a new dataset, OpenConjecture, which contains most conjectures from math papers published on the arXiv since February 2026. We apply frontier LLMs to a sample of 200 of these using a reproducible protocol without problem-specific mathematical guidance. Taken together, the various models produce candidate solutions for roughly 60% of the conjectures, with counterexamples more common than proofs. We then perform a qualitative analysis of solution characteristics such as novelty and connections to existing literature. Our analysis distinguishes proposed conjecture resolutions that expose missing assumptions or recover existing results from those that may contribute new mathematics. Finally, we document the prompts and computational resources required, and release the model responses to support further study.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.