Training Hybrid Quantum Policies with Classical Surrogates for Actor–Critic Reinforcement Learning
Abstract
Reading out a parametrized quantum circuit through finite-shot measurements introduces a discrete stochastic component into hybrid quantum-classical neural networks, making gradient-based training inefficient. In this work, we introduce a surrogate-assisted actor-critic method for training hybrid quantum policies that avoids differentiating the quantum circuit during backpropagation. A classical neural surrogate trained on the circuit's input-output pairs from a replay buffer provides a differentiable substitute for actor updates and critic-target computation, while forward interaction and evaluation retain the quantum circuit's measurements. Our experiments on standard MuJoCo continuous-control tasks demonstrate promising policy learning. Notably, on Humanoid, more runs reach high returns than with learned classical baselines. Moreover, the best average return in our ablation comes from a smaller circuit / surrogate, suggesting that strong performance does not depend on a large network. We also derive conditional upstream gradient-error bounds that separate surrogate fitting from the mean-feature relaxation. Together, the results support surrogate-assisted training as an effective and efficient approach to learning hybrid quantum policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.