Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning
Abstract
Cooperative multi-agent reinforcement learning is well suited to problems with large parameter spaces and exploitable local structure, however, if parameter coupling is strong, a non-stationary environment from the perspective of any individual agent can destabilize learning. We propose using a factored representation of the action space, learned online, to decouple agents. We demonstrate this approach on the tuning of electrostatically defined quantum dot arrays, a leading platform for quantum computing but whose calibration is hindered by cross-talk between gate electrodes. Our framework, **QADAPT**, optimizes voltages using local rewards obtained from individual agents. Through a parameter-sharing strategy, QADAPT also generalizes zero-shot to unseen array sizes, maintaining an approximately constant number of convergence steps. This modular approach provides a scalable route toward rapid calibration of large-scale quantum processors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.