acceptodds
Under review as a conference paper at ICLR 2027

Security-Impact-Validated Traces for Module-Level RTL Vulnerability Detection with Large Language Models

Abstract

Detecting security weaknesses at the register-transfer level (RTL) can reduce costly post-fabrication repairs. Large language models (LLMs) offer automated detection, but require module-level data with reliable security labels. We present a Trace-aware detection approach and a data-construction pipeline combining Common Weakness Enumeration (CWE) scenario identification, structure-adaptive semantic mutation, whole-module syntax checks, and CWE-specific security-impact validation. The pipeline produces 2,444 deduplicated secure–vulnerable pairs across seven open-source platforms and eight hardware CWEs. Aligned views provide code-only inputs (L0) or code with Security Trace annotations (L1): validation-derived code blocks, control conditions, signal relationships, and statically resolvable dependencies, excluding endpoint values that could reveal labels. We compare code-only detection (No Trace), reference evidence (Gold Trace), and model-generated evidence (Self-Generated Trace), using source-lineage-isolated splits and platform holdout to assess detection and generalization. On 492 test modules, four models with source-ordered chunk aggregation achieve 96.54%–97.76% accuracy with Gold Trace; replacing it with generated Trace while retaining the same reference-trained classifiers yields 85.77%–92.68%. These results support the utility of structured security evidence while highlighting the gap between reference and generated Traces.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.