VeriBench-Infinity: Detecting Verification Hallucinations, Valid Proofs of the Wrong Program
Abstract
Formal verification aims to prove that a program satisfies its specification for every valid input, using an interactive theorem prover such as Lean 4. Yet AI agents can produce valid proofs about mistranslated programs or weakened requirements, creating false confidence in software that still contains bugs. We introduce **VeriBench-Infinity**, a benchmark and evaluator that detects these *verification hallucinations* by checking whether AI-generated program translations and specifications preserve the source program's behavior and intended requirements. For each Python program, we independently construct a reference implementation and specification in Lean. We cross-check the reference against Python through execution tests and independent review. Rather than accepting LLM judgment or agreement on finite tests, a separate prover must **formally establish** that the agent and reference agree on observable behavior and requirements for **every valid input** represented by the reference. These comparisons accommodate different data representations and produce machine-checked proofs or concrete counterexamples; incomplete proof searches remain unresolved. We evaluate three agents on 48 program instances pairing correct and buggy versions of competition problems and repository functions. Our analysis exposes translation and specification errors that survive LLM review, while separate case studies identify bugs in CPython's standard library and `intervaltree`. VeriBench-Infinity shifts evaluation from rewarding valid proofs in isolation toward measuring whether AI-generated correctness claims justify trust in the software developers actually use.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.