acceptodds
Under review as a conference paper at ICLR 2027

VeriScale: Adversarial Test-Suite Scaling for Formal Specification Evaluation

Abstract

Formal verification can establish that generated code satisfies a formal specification, but this guarantee is meaningful only when the specification faithfully captures the intended behavior. Existing benchmarks evaluate LLM-generated specifications using concrete test cases, whose sparse coverage can allow flawed specifications to pass and thus systematically overestimate SpecGen capabilities. We present VeriScale, a framework for scaling test suites for formal specification evaluation. VeriScale expands candidate inputs through LLM-based seed generation and type-aware mutation, and synthesizes adversarial implementations that exploit weaknesses in generated specifications to produce challenging incorrect outputs. It further constructs compact test suites through boundary-preserving reduction and mutant-guided selection, reusing the adversarial implementations as mutants. Applied to Verina, VeriScale produces VerinaPlus, an approximately expanded benchmark, and VerinaLite, a compact variant. Across eight state-of-the-art LLMs, VerinaPlus reduces the average Spec Score from 61.44 to 55.09, exposing approximately 12 additional task-level failures per model among 189 tasks. VerinaLite obtains an average score of 56.02, retaining 85.45% of the score reduction with only about 16% of the tests in VerinaPlus. Code and data are provided in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.