acceptodds
Under review as a conference paper at ICLR 2027

Reassessing Software Intermediates in Spec-to-RTL Generation: Accuracy, Cost, and Failure Boundaries

Abstract

Writing correct register-transfer-level (RTL) designs in hardware description languages (HDLs) such as Verilog requires substantial engineering effort. High-level synthesis (HLS) offers a higher-level alternative by compiling behavioral descriptions in languages such as C and C++ into RTL. As modern models become substantially better at direct RTL generation, we ask whether these software relays still justify their added calls and translation boundaries. We study this question by separating the end-to-end pipeline effect of using a software relay from the incremental artifact effect of adding a software view while holding the final hardware rules and authoritative contract fixed. Across three repeats on 156 VerilogEval v2 and 50 RTLLM tasks, direct RTL achieves the highest observed pass rates (88.2% and 58.0%) and lowest token cost among eight evaluated pipelines. In a separate 100-task, three-repeat control, adding frozen Python/C++ views alongside an intact hardware contract yields no observed accuracy gain while increasing final-lowering token use by 18.8%/35.6%, excluding view construction. Controlled ablations and selected failure traces examine intermediate generation, execution, and RTL lowering, complemented by an analytical framework explaining why intermediate test success does not guarantee final RTL correctness. Together, these findings support direct RTL as the reference correctness-cost baseline in the evaluated module-level settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.