MathAdv: beyond proof accuracy in advanced mathematical reasoning
Abstract
A verified proof establishes correctness for a formal statement, but aggregate proof accuracy offers limited insight into the capabilities behind model success or failure. We introduce Mathadv to probe these capabilities across advanced mathematics and test whether proof success persists under equivalent reformulations. The benchmark contains 321 undergraduate- and graduate-level problems across 13 domains, including 298 with Lean 4 statements. Alongside theorem proving, suitable problems have up to three auxiliary tasks: multiple-choice questions probing mathematical knowledge, direct-answer questions assessing informal reasoning without Lean proof construction, and expert-crafted reformulations testing robustness to problem presentation. Our evaluations reveal limitations that proof accuracy alone obscures: models can identify promising mathematical approaches yet fail to construct valid Lean proofs, and their performance varies substantially across subjects and equivalent formulations. Natural-language guidance further exposes model-dependent differences, helping general-purpose LLMs but sometimes hindering proof-specialized models. Together, these findings motivate evaluating not only whether models produce valid proofs, but also the mathematical capabilities underlying their performance and the conditions under which they succeed. The dataset and evaluation scripts are available at https://github.com/mathadv912-design/mathadv-iclr.git.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.