acceptodds
Under review as a conference paper at ICLR 2027

From Rediscovery to Discovery: Verifiable Evaluation on Open Mathematical Problems

Abstract

Recent progress in AI for mathematics has pushed frontier models from competition-style problem solving toward research-level reasoning, yet how to measure verifiable progress on genuinely open mathematical problems remains largely unresolved. Most mathematical benchmarks evaluate reasoning against known answers, while existing open-problem settings do not clearly distinguish recovering established results, making measurable progress, and resolving open questions. We introduce **OpenProblemBench**, a result-verifiable benchmark of 65 mathematical research tasks spanning **L1 Rediscovery**, **L2 Progressive Discovery**, and **L3 Decisive Discovery**. We systematically evaluate frontier reasoning models, long-horizon agents, and discovery systems under this common setting. Our results reveal markedly different capability profiles across the three levels: frontier models can already rediscover multiple recently established results; long-horizon interaction substantially improves feasible-solution coverage and task-native metrics on L2 Progressive Discovery; yet no evaluated configuration reaches L3 Decisive Discovery. To better convert additional test-time computation into mathematical progress, we introduce **MathMap**, a simple yet effective verifier-grounded harness that combines diverse exploration, external verification, and verified selection. MathMap builds a persistent Verified Discovery Tree by expanding parallel tool-using branches, externally verifying candidate solutions, and continuing from the strongest verified checkpoint. Under matched backbone and time budgets, MathMap achieves rank 1 on 14 of 20 L2 tasks among the four same-backbone discovery systems, while using substantially fewer completion tokens and solution submissions. Together, OpenProblemBench and MathMap provide an open foundation for measuring AI progress from rediscovery through progressive discovery to decisive mathematical discovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.