acceptodds
Under review as a conference paper at ICLR 2027

PlaceReasoner-Pro: An 8B Open-Weight Reasoning Model for Frontier-Level VLSI Macro Placement

Abstract

Reasoning-driven agents have substantially advanced VLSI (i.e., chip) macro placement, but their reliance on proprietary frontier models raises severe concerns on confidentiality, specialization, and inference cost. We present PlaceReasoner-Pro, an 8B open-weight vision-language model for macro placement, trained with supervision that moves step by step from cheap geometric checks to the true objective: post-route power, performance, and area (PPA). Because real designs are proprietary and each place-and-route (P&R) run is expensive, we first build a circuit-agnostic synthetic environment that generates realistic floorplans (core outlines, macro sizes, and pin sides) and automatically verifies that a placement is complete, inside the core, and overlap-free; it yields 10K verified, reasoning-annotated placements for supervised fine-tuning. Two-stage reinforcement learning then refines the model: 1) Stage-I uses this verifier as the reward to make placements geometrically legal, raising overlap-free placement from 78.1% to 90.9%; and 2) Stage-II runs every candidate through a full OpenROAD P&R flow and uses the measured PPA as the reward, improving total negative slack (TNS) over Stage-I on 13 of 16 benchmark tasks. On these 16 tasks, PlaceReasoner-Pro achieves timing comparable to Claude OpusĀ 4.8 in the same harness (7 vs. 7 wins on worst negative slack, 7 vs. 9 on TNS) while running entirely on premises. Our results show that compact, domain-specialized small models can approach frontier-level macro placement without relying on proprietary inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.