Transformal: Towards Semantically Faithful Formal Repository Translation in the Wild
Abstract
Large verified developments are scattered across proof assistants with different logics, such as Coq, Isabelle, and Lean. Formal repository translation enables their reuse across systems, but ensuring that translations preserve meaning remains an open problem. Even checking this translation faithfulness is difficult, because it requires relating declarations across two logics. Most existing faithfulness metrics check a statement against a reference or an informal description within a single logic, so they do not directly apply to translation. To bridge this gap, we propose counterexample-based faithfulness (CEF), a reference-free metric that reduces faithfulness checking to counterexample finding. A counterexample consists of source and target inputs that semantically represent the same object but lead to incompatible observations. To make counterexample search more reliable and efficient, we build Transformal, an LLM agent supported by four stages of program-analysis tools. The stages extract dependencies and elaborated signatures from the compilers, derive candidate inputs from branch conditions, and replay each candidate in both proof assistants. To evaluate Transformal, we construct TF-Judge, a benchmark of 297 expert-labeled declaration pairs from 18 real-world Coq and Isabelle repositories. With Claude Opus 5, Transformal reaches 97.67% accuracy, compared with 78.79% for the same model judging directly. The counterexamples of Transformal also provide precise feedback for translation, because each one shows where a translation diverges. On our translation benchmark TF-Translation, this feedback raises the share of GPT-5.6 Sol's declarations that satisfy CEF from 80.25% to 92.48%, four times the gain of compiler feedback. Counterexamples thus make translation faithfulness both measurable and repairable, and Transformal finds them more reliably and efficiently than baselines. Our code and data are available at https://anonymous.4open.science/r/transformal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.