Specat: Assessing and Categorizing Formal Specifications from Natural Language
Abstract
As AI systems generate code at increasing scale, providing strong assurance such as formal verification becomes increasingly important. However, formal verification is only as good as the specification it verifies. In this work, we study formal specification synthesis and evaluation, a standalone but equally important problem. We first formulate spec-evaluation in the real-world setting: assessing a synthesized formal specification against a brief natural-language(NL) description and a few behavioral examples indicating user intent. We introduce the Spec-Evaluation Methodology, which assesses the relations among an NL description, behavioral examples, a candidate formal specification, and formal test artifacts. It interprets verifier or execution outcomes with critic assessments. We further propose SPECAT, an agentic workflow that synthesizes formal specifications and refines them through two complementary loops. Behavioral-oracle refinement uses user-supplied examples and verifier feedback, while counterexample-guided refinement uses an LLM-based Spec Critic to seek description–specification mismatches beyond those examples. Proposed challenges are formally evaluated, while ultimately under-specified or unresolved cases are left for human clarification. We evaluate SPECAT on 704 benchmark-backend units spanning Dafny, Verus, Why3, and Frama-C. Rather than returning a specification alone, SPECAT categorizes each candidate with an evidence-backed disposition: 391 are Endorsed, 77 are Repaired following a machine-supported challenge, 70 retain an unrepaired challenge (Defect), and 86 are deferred because the available evidence is Unadjudicated or Inconclusive. The remaining 80 fail behavioral-oracle refinement. Together, these results show that SPECAT turns specification generation into a testable assurance step: it produces not only a candidate contract, but also machine-grounded evidence and an explicit disposition for that contract.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.