acceptodds
Under review as a conference paper at ICLR 2027

Isomorphism Is Not Enough: An Agentic Audit of Verified Code Generation in Lean

Abstract

Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabilities extend to program verification and how reliably existing benchmarks can measure them, we audit CLEVER, a Lean 4 benchmark for verifiable code generation, using three independent agentic stacks based on Claude Opus 4.6, GPT-5.4, and Kimi K2.5. Our audit finds arguably erroneous reference specifications in 80 of 161 entries, each accepted by the benchmark's authors, distills a taxonomy of benchmark errors, and exposes structural limitations of isomorphism-based scoring, whose signal breaks down when generated and reference specifications formalize different yet defensible readings. In the process, the stacks generate arguably valid specifications for up to 98.8% of entries, certify implementations against sound reference specifications for up to 92.3%, and reach up to 96.2% end-to-end success over entries with sound premises, suggesting that CLEVER no longer stresses modern agentic provers as much as its design intended. Looking ahead, we translate our findings into recommendations for future work, with our error analysis doubling as a repair list for CLEVER and a heuristic checklist for the next generation of benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.