acceptodds
Under review as a conference paper at ICLR 2027

MathAtlas: A Benchmark for Autoformalization in the Wild

Abstract

Prior autoformalization benchmarks are largely focused on olympiad or undergraduate mathematics, while graduate and research-level mathematics remains under-studied. In this paper, we introduce MathAtlas, the first large-scale autoformalization benchmark of in-the-wild graduate-level mathematics, containing theorems, definitions, exercises, examples, and proofs extracted from 103 graduate mathematics textbooks. MathAtlas is enriched with a mathematical dependency graph containing relations, and is the first autoformalization benchmark to include such relations, facilitating evaluation and development of dependency-aware autoformalization systems. Our extensive experiments reveal several gaps in current autoformalization systems. First, we find naive, single-pass systems are insufficient for high-level mathematics. Second, we find feedback-aware formalization loops are an improvement, but primarily only help compilation rate, not semantic faithfulness. Third, we find frontier agentic systems, though effective, are extremely expensive and inefficient, thereby limiting at-scale autoformalization and making efficiency a key challenge. Furthermore, we find system performance degrades substantially with dependency depth. To study performance on “deep” mathematics, we release MA-Hard, a split of 700 entities with the deepest dependency trees. On MA-Hard, our strongest agentic system achieves 64% correctness, but does so inefficiently. We release MathAtlas to the community as a benchmark set for large-scale autoformalization of graduate-level mathematics in the wild. Finally, we introduce MA-Align, a semantic-faithfulness benchmark containing 200 labeled graduate-level examples covering both statements and definitions. Existing metrics perform significantly worse on MA-Align than on prior undergraduate-level benchmarks, indicating a need for further work in this area.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.