acceptodds
Under review as a conference paper at ICLR 2027

Aligning Faithfulness for Theory-Scale Autoformalization

Abstract

Autoformalization (AF) requires more than compilation: formal statements must preserve the meaning of their informal sources, a property known as faithfulness. An emerging frontier is the scalable evaluation of this property, yet most existing methods operate at the statement level and produce whole-item verdicts that do not reveal how individual expressions correspond or which contextual interpretations support the judgment. This omission is consequential because AF is inherently theory-scale: target statements derive their meaning from surrounding mathematical material and unstated conventions. To address these limitations, we reformulate faithfulness evaluation as a context-aware align-then-label problem: identifying phrase-level correspondences between informal and formal statements and assigning a faithfulness label to each under explicitly specified contextual interpretations and assumptions. We also introduce FaithAlign, the first theory-scale faithfulness benchmark comprising 460 items and 4,982 annotated phrase-level links, alongside FaithAlign-Gap, a companion set of 146 items with informal context and assumptions withheld. Each FaithAlign item includes the target statement pair, relevant context, formal dependencies, and, where needed, assumptions specifying the interpretation under evaluation. Across 8 models, GPT-6 Astra achieves the highest defect F1 at 47.8% and identifies every defect in only 4.5% of items containing two or more defects. On FaithAlign-Gap, Astra correctly identifies only 28.2% of undecidable items as such. These results highlight the difficulty of recovering context-dependent semantic correspondences, establishing FaithAlign as a testbed for fine-grained, context-aware faithfulness evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.