FormalMathArena for Reliable Evaluation from Natural Language Problems to Lean Proofs
Abstract
Recent advances in LLM-based theorem proving have largely been evaluated through proof-only benchmarks or separate assessments of formalization and proof quality. Such evaluations neither provide a unified accurate measure of end-to-end theorem-proving capability nor reliably identify the source of failure. To address these limitations, we introduce FormalMathArena, an agentic workflow for evaluating the entire formal theorem-proving process. FormalMathArena combines semantic alignment, proof-assistant verification, and proof-guided failure diagnosis to accurately and separately measure formalization accuracy, proof accuracy, and overall end-to-end accuracy. We further construct a benchmark of 5,427 mathematical problems and validate FormalMathArena on a dedicated expert-annotated subset, confirming its reliability. Applying the validated framework to representative LLMs at scale reveals a gap of up to 21.78% between proof-only and end-to-end performance, identifying accurate and proof-ready formalization as a major bottleneck. We further repurpose FormalMathArena as an action-verification harness, improving formalization accuracy by up to 15.2% and proof accuracy by up to 14.1% on an additional held-out subset. These results establish FormalMathArena as both a reliable evaluation framework and an effective inference-time verifier for end-to-end formal theorem proving. Our code and data will be released once accepted.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.