RegimeQ: Reuse, Sampling, and Candidate Retention in Online QAOA Execution
Abstract
After QAOA parameters have been inherited from a related problem, a finite sampling budget must be divided between further parameter search and solution readout. We propose RegimeQ, a candidate-feedback policy that uses four low-shot probes to select between these actions. Probe costs are included in the window budget, and the decision rule is frozen before evaluation. We evaluate RegimeQ on sequential portfolio optimization problems with 18, 20, and 22 assets, using X-mixer circuits for the main experiments and additional depth and XY-mixer checks. At 16,384 shots per window, the recent eighteen-asset X-mixer evaluation achieves a mean gap of 3.2718% with six calls, compared with 4.4052% and five calls for probe-then-readout, and 3.6035% and 6.48 calls for random allocation. In a separate 2018–2019 evaluation period, the unchanged rule achieves the lowest mean gap among the tested policies at both 18 and 20 assets. The eighteen-asset XY evaluation also shows lower mean gap and fewer calls than random allocation. Matched-sample analysis reveals how retaining probe candidates can reverse aggregate action rankings, while search frequency and duration jointly determine call cost. Together, these results demonstrate favorable quality–call tradeoffs in several settings and identify how parameter inheritance, probe geometry, and candidate retention shape allocation outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.