Toward Fully Unitary Quantum Backpropagation
Abstract
We introduce a fully unitary quantum backpropagation framework that trains deep quantum neural networks without any mid-circuit measurement. The algorithm operates in a hybrid Schr\"odinger–Heisenberg paradigm: the forward pass alternates weight unitaries with QSVT activation layers (polynomial approximations of non-linearities such as sigmoid or tanh, realised as block-encodings), caching both pre-activation states and activation states in ancilla registers; the backward pass propagates the loss operator through both activation and weight layers in the Heisenberg picture and coherently writes gradient states into ancilla registers—entirely without measurement collapse. Three algorithmic enhancements—an operator cache, a barren-plateau detection-and-recovery mechanism, and a quantum Adam optimiser with warmup and cosine decay—are each supported by formal guarantees. Our main theoretical contributions are: (i) an per-iteration gate bound ( the gradient QSVT degree, the activation QSVT degree, the maximum layer depth, the number of layers) obtained via a quantum circuit-wrapping model that propagates the loss operator as a nested block-encoding—exact, with no Pauli-string enumeration or truncation error; (ii) an convergence rate matching classical Adam whenever the QSVT accuracy satisfies ; (iii) a probabilistically sound-and-complete barren-plateau detection guarantee for bounded gradient norm (); and (iv) a near-matching lower bound, establishing near-optimality when is polylogarithmic in . As an alternative propagation strategy we also develop, in app:pauli-mode, an approximate Pauli-propagation mode with a certified truncation-error bound . Empirically, on the MNIST digit-parity benchmark our circuit-wrapping gradient matches a state-of-the-art Mixture-of-Quantum-Experts baseline at every expert count while incurring less quantum-circuit cost per training step; and because this cost is independent of the parameter count, it affords deeper experts that lift test accuracy to —exceeding the baseline's best reported —at a per-step cost reduction the baseline cannot match.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.