acceptodds
Under review as a conference paper at ICLR 2027

MARSA: Joint Learning of Qubit Mapping and Routing with Verifiable Self-Improvement

Abstract

Executing a quantum circuit on a processor with limited connectivity requires choosing an initial qubit mapping and a sequence of routing operations. These decisions are often exposed separately, but they are fundamentally coupled: a mapping is useful only through the routing trajectories it enables, while the best route depends on the mapping from which it starts. We study this coupling directly with **MARSA** (**M**apping **A**nd **R**outing with a **S**emantic **A**utomaton), a single autoregressive policy that generates an initial mapping followed by the complete physical routing trace. This representation creates a structured-generation problem: the legality of every new token depends on the current logical-to-physical mapping, the hardware graph, and the unexecuted circuit. We therefore introduce a *route-state automaton* that maintains these semantics progressively and masks invalid actions during both inference and reinforcement learning rollouts. To learn this long-horizon structured policy efficiently, we initialize from a strong heuristic expert using a two-axis curriculum over circuit width and length, then move beyond imitation through *verifier-guided self-improvement*: MARSA generates its own complete solutions, exact verification evaluates their correctness and SWAP cost, and GRPO reinforces the better generations. On Q20 random circuits, MARSA strictly improves on the training expert for 81.12% of circuits, matches or improves it for 97.15%, and reduces aggregate SWAP count by 16.40%; it also transfers from synthetic training to held-out real quantum algorithm circuits, reducing aggregate SWAP count by 71.35% relative to Qiskit LightSABRE. A separately trained model on a second 2D grid coupling graph also achieves strong performance. Together, these results show that joint mapping and routing can be cast as structured generation, enabling a learned compiler to optimize both and continue improving through its own verified solutions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.