Learning from Good Placements: Memory and Adaptive Self-Imitation for Macro Placement
Abstract
Reinforcement learning for macro placement relies on intermediate placement information to guide hundreds of sequential decisions, whereas downstream quality can only be assessed after completing the full sequence and running a costly global-placement stage that includes standard-cell placement. Moreover, downstream placement quality is not necessarily well correlated with intermediate macro placement objectives, and these relationships can vary across benchmark designs, while the relative importance of these objectives is typically fixed throughout policy optimization. This limits both alignment with the downstream objective and the ability to adapt to the characteristics of individual placement problems. We address this with COMPAS, a Constrained Optimization framework for Macro Placement via Adaptive Self-Imitation, which incorporates sparse downstream placement feedback directly into policy learning. Downstream evaluation is incorporated into COMPAS through two pathways. Global Placement (GP) improving placements update a GP-selected reference memory that guides subsequent search, while selected evaluated trajectories directly supervise the policy through adaptive self-imitation (ASI). During standard policy optimization, PPO is trained using dense wirelength and reference-placement guidance rewards. Periodically, complete macro-placement trajectories are evaluated by the downstream global placer, and selected trajectories are retained for replay, allowing decisions associated with favorable downstream outcomes to directly shape subsequent policy updates. Experiments on eight ICCAD 2015 contest benchmarks show that COMPAS reduces mean GP-HPWL by approximately 8.7% relative to EXPlace, the strongest RL-based baseline in our evaluation, improving seven of eight benchmarks with reductions of up to 22%. Additionally, COMPAS discovers 44% more new GP-HPWL incumbents than the non-ASI variant, with the cumulative gap widening over the optimization horizon. This trend suggests that ASI continues to improve downstream search later in training by replaying trajectories associated with favorable global-placement outcomes. In downstream physical-quality evaluation, COMPAS further achieves 9.6% lower average rWL, 7.9% lower NVP, and a 49% reduction in TNS magnitude relative to EXPlace. Together, these results show that sparse downstream evaluation can both guide subsequent search and provide direct policy supervision through adaptive self-imitation during sequential macro-placement optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.