acceptodds
Under review as a conference paper at ICLR 2027

RISCForge-Bench: Benchmarking the Gap Between Local RTL and Processor-Extension Delivery

Abstract

Module-level RTL benchmarks ask whether an AI system can generate functionally correct hardware. A deployable processor extension must instead preserve one software-visible operation across software, an ISA contract, executable semantics, temporal interfaces, a processor core, workloads, and implementation constraints. We introduce RISCForge-Bench, a family-based executable benchmark spanning 12 independently admitted RISC-V extension families across four workload domains. Each family is one statistical breadth unit; together, they define 48 registered variation axes and 8,403 enumerated sealed-test entries. Across matched model–agent configurations, systems pass 95/108 isolated compute-RTL cells. When systems also author processor-facing RTL, 6/108 cells pass integrated RTL functionality, but none passes protocol stress. End-to-end delivery under full cross-artifact authorship remains 0/108. Same-artifact replay finds 177 locally correct RTL modules and no corresponding delivery. Supplying the admitted implementation restores 32/32 provider-completed deliveries, showing that the evaluator can accept correct artifacts. Responsibility controls place the earliest tested break at processor-facing RTL without claiming to isolate protocol reasoning. Targeted repair recovers nonzero delivery from a near-complete implementation. These results expose a transfer gap between local artifact competence and cross-artifact executable consistency, and establish a versioned target for evaluating long-horizon hardware agents under bounded engineering budgets and auditable scoring.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.