acceptodds
Under review as a conference paper at ICLR 2027

Scaling Laws for Higher-Order Attention

Abstract

Language-model scaling relies on ever more training data, and treats per-token computation as a cost to minimize. As high-quality data runs out, we ask whether increasing per-token computation can improve performance instead. We compare scaling laws of standard and higher-order attention, where each query attends jointly over pairs of keys, making attention cubic rather than quadratic in sequence length. To enable this, we release dense and dilated higher-order-attention Triton kernels with ALiBi3D positional encoding, supporting long-context training and length extrapolation. At matched parameters and data, higher-order attention attains lower loss; standard attention needs – more tokens to match it. The gain stems from a lower irreducible loss ( nats), not a faster data exponent, and carries over to downstream tasks. We then ask whether higher-order attention is more efficient. We find that it depends on the resource: in our regime, higher-order attention wins per parameter and per token, while standard attention wins per FLOP and per byte of KV cache. Because higher-order attention has a lower loss floor, our fitted laws do predict it overtakes standard attention per FLOP and per byte of KV cache beyond our training scale, signaling the promise of higher-order attention at large compute budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.