Structure-Aware Equivalence for Evaluating Autoformalized Statements
Abstract
Evaluating autoformalized mathematical statements requires determining whether a candidate preserves the mathematical claim intended by a trusted reference. Existing methods approach this problem from complementary directions: reference-free learned evaluators estimate semantic agreement but are not constrained to a formal equivalence relation, while reference-based proof methods typically employ whole-statement logical equivalence to provide formal certificates, which is too coarse for translation fidelity and may identify mathematically unrelated statements. We introduce Structure-Aware Equivalence (SAE), a role-preserving formal equivalence relation between syntactic and whole-statement logical equivalence that allows formally justified reformulations while preserving translation-relevant mathematical structure. We operationalize SAE in the Autoformalization Correctness Evaluator (ACE), which provides certified equivalence or inequivalence judgments and reports Inconclusive otherwise. On ProofNetVerif, a benchmark introduced with BEq+, ACE retains substantial equivalent-case coverage while remaining highly conservative in certifying disputed or inequivalent pairs. On targeted diagnostics, BEq+ accepts 94–96% of unrelated but whole-statement logically equivalent pairs, while Mathesis LeanScorer accepts 42% of localized meaning-changing perturbations; ACE accepts none of either set as equivalent.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.