VERDICT: Verifier-Decoupled Inference-Time Cross-Verification for Verilog Generation
Abstract
Test-time scaling for Verilog samples several designs and must return one without a testbench. With the pool's own selector, the return, not the sampling, is where correct designs are lost: on VerilogEval-v2, nine samples from two cheap model families contain a correct design for of problems, more than Claude-Opus-4.8 solves in one shot, yet the pool's cross-verifier returns one for only , and twelve pool-internal selectors all stay below one-shot. We present VERDICT, a verifier-decoupled, testbench-free harness that separates selection from generation. VERDICT freezes a cheap pool so that every selector is scored on the same candidates; makes selection a swappable slot whose default extends ChipMATE's Verilog-Python cross-verification; and, before the slot, runs execution tools that drop candidates which fail to compile or contradict a waveform the specification itself prints, plus a behavior gate that calls the selector only where survivors differ. Filling the slot with Claude-Opus-4.8, which judges but never writes Verilog, recovers percentage points on VerilogEval-v2 and on RTLLM-v2, positive on every seed, bringing the harness to , statistically indistinguishable from a matched frontier one-shot () that writes the design itself. On VerilogEval-v2 the tools add points to the unchanged cross-verifier without losing a problem; waveform replay covers the of specifications that print one and was calibrated on them. Run over the frontier model's own samples, the same harness reaches against one-shot's (not significant). VERDICT's accounting also explains when a harness pays off: a pool gains over one-shot exactly when its margin exceeds its selection loss. VERDICT thus turns "sample more" into three quantities to report: coverage, selection loss, and cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.