TrustMe: A Benchmark for LLM-Aided Hardware Design Code Auditing
Abstract
The advancement of Large Language Models (LLMs) has accelerated the adoption of vibe coding in chip design, catalyzing the emergence of LLM-Aided Design (LAD) as an agile development paradigm now widely deployed in practical chip development workflows. This work focuses on a critical task in the chip supply chain: design code security auditing, with a primary emphasis on hardware Trojan detection. Identifying security threats at the early stages of the chip lifecycle is of paramount importance. Existing LAD research has predominantly centered on code generation, verification script synthesis, and debug completion, with limited attention paid to code security alignment and hardware design security awareness. This gap may lead LLMs to generate or overlook hazardous design code, thereby contaminating downstream segments of the supply chain. To bridge this gap, we present TrustMe, a comprehensive benchmark for hardware code Trojan security auditing. TrustMe comprises 816 tasks sampled from real-world projects, with over 1560 complete and stealthy Trojans implanted. It spans evaluation scenarios from module-level to System-on-Chip (SoC) level, assessing full-stack hardware code security monitoring capabilities, including knowledge-based question answering, vulnerability localization, and attack type analysis. The benchmark supports simulation testing across diverse vibe coding environments, including both agentic and pure-LLM setups. We evaluated 10 mainstream LLMs and 6 agentic tools on TrustMe. Results show that even the best-performing combination achieves a hardware Trojan identification success rate below 50%, underscoring the urgent need to prioritize security considerations in LAD. Furthermore, based on our evaluation findings, we propose a Skill-Powered Trojan detection and identification framework. Operating in a self-iterative improvement mode, this framework achieves an average 12.6% performance boost without modifying existing models or agentic tools. We hope this work will encourage LLM-aided chip design from productive into genuinely trustworthy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.