QuARC: A Benchmark for LLM-Generated Quantum Circuits and an Audit of Its Acceptance Region
Abstract
Scoring an LLM’s code against a reference can reject correct solutions. Tolerating equivalences can remove such false rejections, yet each tolerated transformation can widen what the verifier accepts. QuARC brings this tradeoff to quantum circuits, scoring 86 tasks by membership in an explicit acceptance region rather than distance to a reference. The region admits any circuit meeting the task’s predicate up to global phase, gate decomposition, qubit relabeling, and ancillas that start and end in the all-zero state. Relabeling, however, applies only where a task fixes no layout. An audit on probes and baseline generations measures what each design choice prevents and admits. A check at the all-zero input alone, for instance, accepts 42 generations wrong under their prompts, none of which QuARC’s verifier accepts. Relabeling decides 76 acceptances, 8 of them wrong under the prompt’s labels, and no false rejection that it prevents is established. Restricting relabeling to prompts naming no qubit roles or order would remove these wrong acceptances. Separately, a disguised copy of the canonical output passes 16 of the 70 tasks without free parameters, all among the 17 judged at a single input. Across the evaluated models, pass rates span 55.8% to 91.5%. We release the corpus, verifier, audit scripts, and baseline generations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.