acceptodds
Under review as a conference paper at ICLR 2027

Understanding Disagreements in Lean Formalization Evaluation

Abstract

Evaluating the faithfulness of a Lean formalization requires interpreting both its natural-language source and its formal statement. LLM judges, human reviewers, and published labels can disagree because of translation defects, mistaken judgments, or differences in interpretation and grading criteria. We investigate these disagreements using three LLM judges on CriticLeanBench, MA-Align, and 208 formalizations with existing human ratings. An author audit of 75 CriticLeanBench items includes all 51 cases where the judges unanimously contradict the label: the audit supports the judges in 44 cases, the label in 2, and leaves 5 unresolved. Source comparisons and targeted Lean checks reveal missed translation defects, ambiguous conventions, and incorrect model objections. LLM judgments also help identify formalizations on which human reviewers disagree, without using human ratings to select them. Together, these findings show how LLM judgments can guide review, while source comparisons and targeted checks distinguish translation defects from differences in interpretation and grading criteria.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.