TREAD: A Transferable RTL Optimizer via Rewrite Trajectories and Synthesis-Grounded Supervision
Abstract
Improving RTL before synthesis is an important way to reduce hardware cost, but effective rewrites depend on both the design structure and the downstream synthesis flow. Manually exploring this space or rediscovering useful transformations for each design is expensive, motivating methods that can learn and transfer optimization experience across RTL designs. We develop TREAD, a transferable RTL optimizer that learns such experience from rewrite trajectories. Imitation learning first provides an executable policy over legal rule–location actions, while offline reinforcement learning captures long-horizon rewrite effects from complete trajectories and delayed returns. To overcome the limited action coverage of observed trajectories, we generate alternative legal rewrites from the same source states and evaluate these unselected actions under a common Design Compiler flow, providing synthesis-grounded preferences for value and policy learning. At inference, TREAD performs bounded residual search beyond the imitation endpoint and accepts a learned residual only when its predicted utility exceeds a positive margin. Across seven RTL benchmark categories under a common technology library and synthesis flow, TREAD achieves the best QoR among the evaluated RTL optimizers under the reported synthesis flow, including representative e-graph-based and LLM-based systems. These results suggest that combining executable behavior, long-horizon returns, and same-state synthesis feedback enables optimization experience to transfer across RTL designs and improve RTL optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.