acceptodds
Under review as a conference paper at ICLR 2027

QuOOD: Benchmarking Out-of-Distribution Property Prediction on Quantum Circuits

Abstract

Machine-learned surrogates that predict the noisy outputs of quantum circuits or mitigate their errors learn from classical simulation, whose cost grows exponentially with the number of qubits for general circuits. They are trained on circuits small or structured enough to simulate and deployed on circuits that are wider, deeper, from new families, or run on new devices. Distribution shift is therefore the normal deployment condition. Existing datasets test these surrogates in distribution or one shift at a time. We introduce QuOOD, a benchmark that makes performance under these shifts measurable: 99,386 circuits and 394,684 instances, split into an i.i.d. control and four shifts in depth, width, circuit family, and device noise, with every test regime kept within the range where simulation still gives exact or error-bounded labels. Because many target values are close to zero, a low average error can hide both a model that predicts nothing and one that invents structure. QuOOD checks its farthest test regimes for signal and for a clear margin over predicting zero, and adds a control stratum of qubit pairs whose correlations the circuit's structure certifies as near zero. Across thirteen predictors, tree ensembles and neural models swap places: trees over engineered features lead every neural model at 2–3× the training width, where each target depends only on a bounded set of gates, and trail them on the held-out family. Graph networks fall further behind as width grows, even our noise-conditioned transformer with its width inputs removed. On the controls, the transformers (ours and GraphGPS) emit more spurious correlation than their measured input carries, whereas the trees shrink it. Our transformer leads on held-out calibrated devices through per-block noise conditioning. On the held-out family, validation error stops tracking the test ranking (Spearman ρ = 0.00, against 0.99 in distribution), so models meant for shifted deployment must be tested under shift, which QuOOD enables. We plan to release the dataset, splits, labels, and reference runs upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.