TurboRTL: Rigorous Evaluation of LLMs and Agents for RTL Optimization
Abstract
Although large language models (LLMs) have demonstrated growing capabilities in Verilog code generation, existing research largely prioritizes syntactical and functional correctness, overlooking the critical hardware design goal of optimizing circuit efficiency. Consequently, the field lacks a well-established benchmark to rigorously assess the ability of LLMs to optimize register-transfer level (RTL) designs. To this end, we introduce **TurboRTL**, a comprehensive Verilog code optimization benchmark constructed via an automated pipeline that utilizes equivalence graph (e-graph) rewriting to systematically identify designs with significant optimization potential. Furthermore, reflecting the reality that hardware engineers optimize designs through iterative refinement rather than single-shot attempts, our evaluation framework extends beyond individual models to incorporate two tool-augmented interaction paradigms, specifically population-based test-time search and experience-driven agentic workflows. We demonstrate that while individual models show initial promise, these advanced agentic strategies are essential for tackling the complexities of this domain. By providing both a high-quality benchmark and a diverse evaluation pipeline, this work establishes a critical foundation for advancing automated RTL optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.