RTL-BenchLS: A Large-Scale Benchmark for Agentic RTL Reasoning and Implementation
Abstract
Agentic frameworks are advancing rapidly, driving growing interest in agents for automated hardware design. Implementing hardware in RTL (i.e., RTL generation) is a key task in VLSI circuit design. However, existing RTL benchmarks are too easy for cutting-edge agentic frameworks. Specifically, we observe three gaps in the RTL benchmarks. (1) Limited task scope. Benchmarks are limited in module-level scope. Such module-level scope is insufficient to evaluate agentic frameworks. The EDA community urgently requires a qualified project-level RTL benchmark. (2) Narrow task type. Most benchmarks only focus on a narrow task scope. A single narrow task type is not enough to thoroughly evaluate agents on complicated RTL tasks. (3) Small benchmark scale and design size. Existing benchmarks are usually in small scale with small design size. To address these limitations, we introduce RTL-BenchLS, a large-scale benchmark of 17K+ cases with 12 tasks. We contribute in three aspects. (1) New task scope. We expand the scope from module level to project level with realistic project-level designs. Experiments show that our module-level cases are sufficiently challenging, and project-level tasks expose obvious limitations of current agentic frameworks. (2) Expanded task type. We cover three important types to evaluate agents' abilities from different perspectives, including generation, completion, and debugging. (3) Enlarged benchmark scale and design size. We collect a significantly larger scale of cases covering larger design size. To further increase the benchmark scale, we define a new task protocol named round-trip protocol, which eliminates the need to manually craft qualified specifications and testbenches. As a result, we can significantly expand the benchmark scale using collected RTL designs. The round-trip tasks are also significantly challenging for the current LLMs. This new benchmark, at unprecedented scale, will help evaluate academic work and industrial solutions and improve LLM/agent solutions for hardware RTL design. Our benchmark will be released at https://anonymous.4open.science/r/RTL-BenchLS-A0AB.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.