acceptodds
Under review as a conference paper at ICLR 2027

SILICA-BENCH: EVALUATING LONG-HORIZON LLM AGENTS THROUGH VERIFIED RTL-TO-GDS IMPROVEMENT

Abstract

Long-horizon LLM agents must turn a sequence of decisions into a reliable result: interpret imperfect feedback, preserve requirements, and divide a finite budget between exploration and validation. We introduce , which evaluates these abilities through physical improvement of digital circuits with a commercial RTL-to-GDS flow. Four properties make this a useful agent task: an open-ended, continuously graded power, performance, and area (PPA) objective with a fixed behavioral target; hard functional requirements; costly staged evidence over coupled RTL and flow choices; and 36 designs across eight domains and multiple scales. Agents receive reference RTL and a reference flow, work within a declared budget, and register what they deliver. Protected assessment binds RTL and post-route netlist replay to the nominated implementation, using local reconstruction or identity-matched reuse of its physical outputs. The protocol separates validity from improvement, early evidence from final quality, and discovery from delivery. A paired study of six model–harness systems and three edit authorities covers 240 nominations on 30 designs. Every configuration achieves at least 93% validity, yet no Joint system delivers a changed, verified gain above 5% on more than one third of tasks, and two Joint systems fall below unit failure-inclusive quality. Eleven of twelve valid regressions larger than 5% use the local reconstruction path. These results distinguish successful execution from reliable improvement, and make delivery of a verifiable engineering artifact an observable target for long-horizon agent evaluation. We will provide the complete dataset and test environment via an anonymous GitHub link during the rebuttal period.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.