Isomorphism Is Not Enough: An Agentic Audit of Verified Code Generation in Lean
Abstract
Agentic systems have recently emerged as state-of-the-art approaches for automated theorem proving in formal mathematics. To assess how far these capabilities extend to program verification and how reliably existing benchmarks can measure them, we audit CLEVER, a Lean 4 benchmark for verifiable code generation, using three independent agentic stacks based on Claude Opus 4.6, GPT-5.4, and Kimi K2.5. Our audit finds arguably erroneous reference specifications in 80 of 161 entries, each accepted by the benchmark's authors, distills a taxonomy of benchmark errors, and exposes structural limitations of isomorphism-based scoring, whose signal breaks down when generated and reference specifications formalize different yet defensible readings. In the process, the stacks generate arguably valid specifications for up to 98.8% of entries, certify implementations against sound reference specifications for up to 92.3%, and reach up to 96.2% end-to-end success over entries with sound premises, suggesting that CLEVER no longer stresses modern agentic provers as much as its design intended. Looking ahead, we translate our findings into recommendations for future work, with our error analysis doubling as a repair list for CLEVER and a heuristic checklist for the next generation of benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.