CircuitPilot: Learning Circuit Optimization through Imitation and Self-Play
Abstract
Optimizing a circuit means searching a large space of equivalent circuits for a cheaper one, and a bounded optimizer must choose which rewriting and search procedures to run. We propose CircuitPilot, which learns these choices in two stages: a policy first imitates the optimal first actions found by exhaustive short-horizon search, then improves through single-player self-play on targets from policy-guided tree search. We apply one formulation to polynomial, Boolean, and quantum circuits, with a separate policy per domain; every returned circuit is checked for equivalence to its input. On 161 three-qubit development circuits, the quantum policy matches or improves on the cost of a step-limited Quasar configuration on 152 (70 strictly better). An imitation-only polynomial policy obtains lower total cost than each of three fixed schedules, though not lower than their per-input best, and 98 of 100 IWLS 2026 Boolean benchmarks can be encoded as model inputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.