Effective Dimension Governs Generalization in Quantum Kernel Vision Models
Abstract
Recent quantum vision models—quantum vision transformers and quantum convolutional networks—report two striking but unexplained empirical phenomena: (i) ansatze with more, or more uniformly distributed, entanglement generalize better, and (ii) injecting quantum noise can improve test accuracy rather than degrade it. These observations are currently treated as curiosities, discovered by grid search and explained, if at all, by hand. We show that both are manifestations of a single, measurable quantity: the effective dimension of the (noise-shaped) quantum feature kernel. Working primarily with quantum-kernel vision models—a quantum feature map read out by a kernel classifier—we give a spectral account in which entanglement structure and quantum noise are two knobs that move ; in an overfitting regime, contracting acts as ridge-like regularization. We analyze the mechanism: an exact decomposition of the depolarized kernel with , a contraction result (and its boundary) for amplitude damping, a kernel-machine capacity bound, and a capacity/alignment risk decomposition; the monotone contraction operative in our entangled experiments is verified empirically, not proven in general. Our informative empirical finding is that test accuracy collapses onto a single function of across distinct entangling ansatze—compressing different spectral shapes onto one curve ( over seeds). Along the one-parameter depolarizing family the collapse is instead exact by construction; we use it only to confirm the kernel decomposition to machine precision and at up to qubits, not as evidence for . Amplitude damping contracts and lifts test accuracy by up to along an inverted-U sweet spot; the effect's sign flips between the over- and under-fitting regimes; noise injection matches an explicit spectral-filtering frontier (so it is not a weak substitute for hyperparameter tuning); and the phenomenon persists in trained QViT- and QCNN-like models. Entanglement plays the complementary role of a precondition—it supplies the feature-space alignment without which the law does not hold. Our results organize two reported anecdotes into a single measurable principle for designing quantum-vision models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.