Transformers as Random-Lag Autoregressive Models
Abstract
In this paper, we consider a probabilistic generative model approach to explain the Transformer architecture and leads to improved new architectures. We make a minor modification to the classical autoregressive model by replacing its fixed unit lag with a random lag, and prove that the exact posterior mean under the resulting random-lag autoregressive model (RLA) is precisely causal Softmax attention. As a conditional expectation, this attention operator is the unique mean-square-optimal estimator of the latent value readout given the observed prefix. In RLA, all the standard transformer components are endowed with a generative role, such as RoPE, ALiBi positional encodings, and query-key normalization (QK-Norm). By replacing the Gaussian assumption of RLA with a linear directional family on the sphere and again computing the exact posterior mean, we obtain a \em spherical linear attention (SLA). Unlike existing linear attention mechanisms and similar to softmax attention, SLA has normalized and nonnegative attention weights. The experiments show the competitive performance of SLA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.