acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking the Quantum Leap: Auditing LLM Evaluation for Quantum Code Generation

Abstract

Large language models (LLMs) are increasingly capable of generating quantum programs, but reliable evaluation remains difficult: benchmarks differ in task composition, scoring procedures, and failure handling, while leaderboard rankings can obscure uncertainty and evaluator defects. We evaluate 15 hosted frontier, open-weight, quantum-specialized, and retrieval-augmented models across five primary quantum code-generation suites containing 530 tasks, and 10 models on QuantumBenchEval, a previously unreleased collection of 96 tasks across six quantum-computing topics. We report task-level bootstrap confidence intervals, multiplicity-adjusted paired comparisons, and audits of evaluator behavior, output validity, and shared failures. QiskitHumanEval-Standard is approaching saturation for the strongest models (98.7% pass@1), whereas its Hard variant retains substantial headroom (63.8%); across the recorded primary-suite outcomes, none of the five top-two comparisons remains significant after multiple-comparison correction. Evaluation design can have an even larger effect: correcting undisclosed exact-value checks in QuantumBenchEval's T1 topic changes two models' measured pass@1 from 0% to 94.1% without changing their outputs. Our audit also identifies 17 QuanBench-117 tasks with shared failures associated with benchmark-owned imports, while controlled probes of the current evaluator reveal unexecuted unit tests, leakage of canonical functions into candidate execution, and reference imports that can block otherwise valid candidates. These defects can affect both apparent failures and apparent passes, requiring historical QuanBench outcomes to be revalidated under an isolated evaluator. Together, these results show that quantum code-generation progress cannot be characterized by leaderboard scores alone. We release the evaluation harness, per-sample outcomes, and audit artifacts, and advocate protocols that validate scoring and test execution, isolate candidate and reference code, report uncertainty and output validity, and respect sampling assumptions when estimating pass@.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.