acceptodds
Under review as a conference paper at ICLR 2027

MathDB: A Dynamic Evaluation Platform for Frontier AI in Research Mathematics at Scale

Abstract

As AI systems advance at mathematical reasoning, static benchmarks saturate quickly, and repeatedly constructing expert-curated replacements is costly and hard to scale. We introduce MathDB, a continuously updated, structured database of more than 80,000 research-level mathematical problems with literature provenance, domain labels, status histories, verification evidence, and a knowledge graph of mathematical relationships. MathDB grows through automated literature ingestion and community contributions, and its structured metadata supports the construction and later refresh of benchmark releases without resourcing an entirely new problem set from scratch. As a first demonstration of this framework, we construct MathDB-Bench from 100 tasks adapted from 981 resolved MathDB records, spanning five mathematical domains and four reasoning-demand dimensions: structural, symbolic, quantitative/computational, and geometric/spatial reasoning. The tasks use deterministic checkers reviewed by mathematicians for fidelity and resistance to guessing. Evaluating four frontier and open-weight models under a requested 15,000-token completion budget, pass@1 ranges from to . Token-budget truncation, rather than incorrect completed answers, dominates failures: all four models are correct – of the time when they produce a checkable answer. A targeted 30,000-token probe on 10 tasks that were truncated for every model and run at 15k further shows substantial budget sensitivity: GPT-6 Astra solves 6 of 10, while Kimi K3 solves 1 and Qwen3.8 Max solves none. We release benchmark tasks, prompts, deterministic checkers, reasoning-demand annotations, and raw model transcripts, while MathDB serves as the durable contribution: a growing, community-editable resource from which new, unsaturated evaluations can be constructed as existing benchmarks saturate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.