acceptodds
Under review as a conference paper at ICLR 2027

Policy RBMs on Superconducting Quantum Computers

Abstract

Restricted Boltzmann machines (RBMs) are Ising models, and a quantum processor that holds the Gibbs state of an Ising model returns a joint sample of all visible and hidden units of the RBM in a single measurement, without Markov chains or burn-in. Two obstacles stand in the way: the Boltzmann factor is not unitary, so the state is hard to prepare, and superconducting qubits couple to few neighbours, so an RBM whose graph does not match the hardware costs circuit depth that the quantum processor does not have. We treat both as a representation-learning problem in which the quantum processor defines constraints on the model space. A neural network policy for power-grid control, trained by AlphaZero-style self-play, is distilled into a conditional RBM whose 118 units and 210 weights occupy every functioning qubit and coupler of a 120-qubit IBM Nighthawk r2 quantum processor, whose action code follows the measured qubit noise and whose binary input state code is learned with the model. A quantum imaginary-time evolution prepares its conditional Gibbs state. A series of experiments isolates the effect of each technique. The quantum-processor policy agrees with the teacher’s preferred action in about half of the held-out states, against 27 % for the state-independent prior, and lies within a few hundredths of a nat of its exact read-out, whereas QAOA-type circuits fall far short of it. To our knowledge this is the largest RBM sampling experiment so far on a gate-based superconducting quantum processor with a read-out that recovers the trained model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.