The Cost of Quantum Features in Reacher Control: A Resource-Accounted Simulator Study
Abstract
Small quantum circuits have few parameters, but a robot controller may need to run one every time it chooses an action. We study the resulting cost on Gymnasium MuJoCo Reacher-v5. A two-qubit feature map with finite-shot sampling (Q1) and a compact classical control share the same observations, action bounds, PPO schedule, and 50,000 training interactions. We evaluate each of 20 training seeds on the same 100 cases. Q1 finishes 0.00569 m closer to the target on average. Whether it meets the predeclared 0.010 m non-inferiority margin depends on how uncertainty is estimated: the one-sided bootstrap bound is +0.00946 m, while the Student-t sensitivity bound is +0.01074 m. Success remained low: 9.4% for the compact control, 10.3% for Q1, and 11.3% for SB3. Q1 used 57.6 million shots per seed and took 0.915 ms per action at the median, about four times the compact policy’s latency. A later 500k-step comparison retained the full state alongside the features. There, Q1 finished 0.02721 m farther from the target over 18 completed pairs (95% interval [+0.01840, +0.03647] m), and two Q1 runs timed out. The follow-up failed its frozen zero-failure rule and remains exploratory. Q1 success was 16.2% on completed seeds, versus 26.8% for sine. A separate 20-seed SB3 screen improved return on every seed, yet two missed the 10% success floor. The evidence therefore concerns the cost of fixed representations in controllers with limited task success. It does not establish practical robot competence or predict performance on a quantum processor.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.