acceptodds
Under review as a conference paper at ICLR 2027

CircuitPilot: Learning Circuit Optimization through Imitation and Self-Play

Abstract

Optimizing a circuit means searching a large space of equivalent circuits for a cheaper one, and a bounded optimizer must choose which rewriting and search procedures to run. We propose CircuitPilot, which learns these choices in two stages: a policy first imitates the optimal first actions found by exhaustive short-horizon search, then improves through single-player self-play on targets from policy-guided tree search. We apply one formulation to polynomial, Boolean, and quantum circuits, with a separate policy per domain; every returned circuit is checked for equivalence to its input. On 161 three-qubit development circuits, the quantum policy matches or improves on the cost of a step-limited Quasar configuration on 152 (70 strictly better). An imitation-only polynomial policy obtains lower total cost than each of three fixed schedules, though not lower than their per-input best, and 98 of 100 IWLS 2026 Boolean benchmarks can be encoded as model inputs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.