acceptodds
Under review as a conference paper at ICLR 2027

Higher-Order Token Interactions via Quantum Attention

Abstract

A single self-attention layer scores token pairs and aggregates them linearly over values, and known lower bounds put specific order- tasks out of reach for one such layer at budget . We introduce Quantum Higher-Order Attention (QHA), a shallow, hardware-realizable quantum attention head that — via data re-uploading and an all-to-all non-Clifford entangler — synthesizes order- token interactions inside the circuit and exposes them on a local single-qubit read-out. We prove two complementary results, deliberately not billed as a separation theorem because they concern different targets: no budget-limited attention layer computes the order- matching family, while a circuit in the QHA gate set realizes every order- monomial at depth ; and a barren-plateau-free trainability guarantee for the shallow local-design variant — the all-to-all variant we benchmark does show exponentially decaying gradients. Empirically we tune the classical head's learning rate, width and training length to maximize its own accuracy, an advantage we do not grant QHA, and find the gap is graded in : a tuned attention layer learns order-3 parity (), so no separation holds there; from order 4 QHA ( parameters) leads against at ; and at order 6 every attention-paradigm model we test — explicit high-order attention and a larger two-layer Transformer included — sits at chance, the only classical model in our suite that solves it being a larger MLP. A trained order-3 head runs on IBM Heron at accuracy through qubits under a disclosed shot budget, and QHA is the most compact detector on genetic epistasis, learning-parity-with-noise and graph triangle detection. Our claim is an inductive-bias advantage and an empirical learning separation against classical attention, not an exponential speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.