acceptodds
Under review as a conference paper at ICLR 2027

DiffRTL: Differential Reward Models for Synthesis-Efficient Reinforcement Learning of RTL Optimizers

Abstract

Training-time reinforcement learning is the standard route to endowing open-source language models with register-transfer-level (RTL) design ability, but rewarding power, performance, and area (PPA) requires logic synthesis per design, orders of magnitude slower than any other check in the loop. This synthesis bottleneck caps affordable training, and learned predictors of absolute PPA fall short as substitutes, since typical errors exceed the effect sizes of most optimizations. We observe that optimization consumes only the relative effect of a change, which is far easier to learn because paired comparisons cancel systematic errors. DiffRTL trains a differential reward model on verified transformation pairs manufactured by a paired data engine and uses a fast forward pass of the model as the PPA reward, corrected during training by periodic real-synthesis audits. Every reported number remains grounded in real simulation, formal equivalence checking, and synthesis. Spending 640 synthesis calls on audits alone, DiffRTL beats synthesis-rewarded training at six times the budget by 6.8 eff@1 points on Pluto, with matching gains on RTLLMv2, at a small fraction of the tool time. The differential formulation clearly surpasses an absolute predictor trained on identical data. Audits measure reward over-optimization exactly and suppress it throughout training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.