acceptodds
Under review as a conference paper at ICLR 2027

When Do RTL Verifiers Lie? Measuring Proxy-Reward Overoptimization in LLM Hardware Design

Abstract

Reinforcement learning with verifiable rewards (RLVR) and inference-time search are increasingly used to improve generated register-transfer-level (RTL) designs. While these techniques require access to rigorous test scripts, high-quality RTL data, and verification infrastructure, these resources are scarce and often proprietary. Recent pipelines therefore rely on capable teacher models to generate synthetic RTL and automated testbenches. However, optimizing against these scalable verification proxies risks Goodhart effects: visible reward may improve without a corresponding gain in true functional correctness. This risk is especially acute because RTL correctness spans input sequences and clock cycles, whereas teacher-generated testbenches, randomized differential simulation, and bounded model checking (BMC) provide scalable but finite checks, with the fidelity of simulation and bounded checking depending on verification effort. We compare four testbench conditions generated by Grok 4.6 and DeepSeek V4 Flash from either the specification or golden RTL, alongside differential simulation and BMC. A hidden formal-equivalence Oracle evaluates candidate RTL against its golden reference. The study covers 100 tasks, eight Qwen3 configurations, 50,948 candidates, Best-of- search, multi-turn repair, and two-epoch RL. Qwen3-32B reasoning reaches 82% Oracle Pass@64, but teacher-generated testbenches select proved designs on only 61–66% of tasks despite visible success rates of up to 93%; tuned BMC and differential simulation reach 78% and 76%, respectively. Increasing verification effort narrows this gap, while weak direct-mode repair feedback can continue improving visible success after hidden progress stalls. In the Qwen3-4B RL run, the proxy–Oracle gap widens and rewarded formal counterexamples become more frequent without a clear improvement in Oracle proof rate; the 8B run shows a persistent proxy–Oracle discrepancy over the same training budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.