acceptodds
Under review as a conference paper at ICLR 2027

Entanglement as a Structural Complexity Axis: A PAC-Bayesian View of Generalization in Quantum Policies and Value Functions

Abstract

Parameterized quantum circuits (PQCs) are increasingly used as policies and value functions in quantum reinforcement learning (QRL), yet almost all evaluations report only average return, leaving open the question of when and why a quantum policy generalizes. We give a PAC-Bayesian answer. We derive a generalization bound for stochastic quantum policies whose complexity term is controlled not by the raw number of circuit parameters, but by the log-determinant complexity of the Fisher geometry induced by the circuit—a quantity that entangling connectivity inflates by expanding the readout's causal cone. Empirically, in a controlled setting that fixes the number and layout of trainable rotations and varies only the entangling connectivity, this acts as an independent axis of complexity: at a fixed parameter count the train-to-test generalization gap grows with the circuit's Fisher log-determinant complexity, whereas raw parameter count is the weakest predictor of the gap. The bound acts as a ranking certificate: at fixed parameter count it consistently distinguishes the non-entangled circuit from entangled alternatives, something a parameter-counting bound cannot do at all. We confirm the mechanism across settings—supervised classification (– qubits, synthetic and real data), a reward-only quantum contextual bandit, and multi-step value-function generalization—where entangled circuits generalize worse than non-entangled ones of identical parameter count and the gap shrinks with sample size as predicted. Our strongest evidence is in these low-variance decision models; in genuine end-to-end multi-step policy learning the standard entangler's effect is statistically significant but return variance leaves the full ordering only partially resolved, so we frame the paper as a controlled study of the entanglement–generalization trade-off in quantum decision models rather than a solved account of end-to-end quantum RL. A partial-correlation analysis shows screens off the entangling pattern (), consistent with readout causal-cone expansion as the operative mechanism and with the Meyer–Wallach entangling power as an empirical correlate rather than the driving variable; controls (matched training accuracy, alternative readout and optimizer) rule out optimization confounders, and the effect survives execution on an IBM Heron processor under real noise. Our results reframe quantum-policy design around an entanglement–generalization trade-off rather than expressivity alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.