acceptodds
Under review as a conference paper at ICLR 2027

HARVEST: Hierarchically Adaptive RTL Verification via Structured Episode Trajectories

Abstract

Automatic testbench generation with large language models can reduce the cost of hardware verification, but a testbench that compiles and passes its own checks may still miss critical bugs. This creates a fundamental challenge for automated verification: the signal used to decide whether a generated testbench is adequate cannot simply be the agent's own report of success. We introduce HARVEST, a framework for adaptive testbench generation that routes designs among heterogeneous verification agents and uses an external adequacy signal}, mutation kill rate, to identify insufficient testbenches and trigger complementary verification. A contextual bandit selects among three agents targeting structural, semantic, and adversarial coverage. When the selected testbench has low external adequacy, HARVEST escalates to a second agent andcomposes the resulting testbenches into a single suite, preserving complementary tests rather than discarding the first attempt. On 100 open-source Verilog designs, we find that 13% of testbenches that pass their own checks nevertheless kill fewer than one-third of injected bugs, revealing a substantial gap between agent-reported success and actual verification adequacy. Routing to a single agent achieves 91.7% of the oracle quality, while adequacy-gated escalation further improves quality at modest additional cost. Upon acceptance we will release 300 structured trajectories spanning the 100 designs as a benchmark for adaptive verification under external adequacy signals. Together, these results show that external adequacy can provide a reliable control signal for detecting false-confidence testbenches and coordinating complementary verification agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.