SpecGuard: What Should LLMs See to Generate Assertions That Catch RTL Bugs?
Abstract
Hardware bugs that escape the rigorous verification and validation process are costly to fix, often requiring recalls, software patches that mask the error rather than fix it, and similar remedial actions. This motivates rigorous verification of each implementation against its intended behavior. Assertion-based verification captures this intent using SystemVerilog Assertions (SVA), which can be used in both formal and simulation-based verification. However, composing SVAs remains largely manual and time-consuming, causing a bottleneck in time-to-market. LLMs show promise in automating this by translating a design’s natural-language specification into SVA without manual formalization. In practice, SVA generation through LLM is impeded by restating shown implementation behavior, placing claims in the wrong clock cycle, and naming signals outside the checked module. Recent studies show SVAs generated by frontier models detect only 5.1–11.7% of injected bugs. In this paper, we analyze assertion suites generated by existing LLM-based SVA generation methods. We evaluate them against injected bugs and identify four failure modes, along with two attributes that place the resulting assertions beyond repair after generation, rendering them ineffective for actual bug detection. These failure modes and attributes are not captured by the compile and proof rates typically used to evaluate generators. While a stronger model improves these rates, its generated assertions do not necessarily detect more bugs. We subsequently develop SpecGuard, an SVA generation framework that addresses each failure by separating what an assertion must draw from a model versus what can be fixed with design’s structure. SpecGuard was evaluated on the 83-design AssertLLM2 benchmark, using a local 14.8-B model with int4 quantization. Its assertions detected up to 66% of the AssertLLM2 bugs in a single run, compared with 11.7% for the best-performing of six frontier models. Furthermore, on more complex hierarchical designs from NVDLA and Google Coral NPU, it achieves up to a 51% bug-detection rate, compared with 14% for existing state-of-the-art assertion-generation techniques.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.