acceptodds
Under review as a conference paper at ICLR 2027

Self-Attention as Connection Propagation: An Exact Geometric Operator Representation

Abstract

Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood. We view a token sequence as a *vector field over the token-position graph* and identify attention as a *connection walk*: messages are aggregated by a nonnegative walk matrix while being *transported* along each edge by a learned linear map. Within this framework, we derive an exact representation of *single-head attention (SHA)* through a connection propagation step with *constant* transport, and of *multi-head attention (MHA)* through a *single edge-dependent connection walk* whose effective transport is an attention-gated mixture of headwise transports. We further clarify the conditions under which the corresponding operator reduces to a *random-walk connection Laplacian*, highlighting the roles of stochasticity, reversibility, and metric-compatible transports. Empirically, we examine trained Transformers across scales and structures (encoder/decoder). Layerwise diagnostics reveal low-drift middle-layer routing regimes and pronounced diagonal structure in averaged transport Gram matrices, while a separate selected-edge spectral analysis reveals concentrated transport energy with unequal directional gains. Overall, the paper provides a precise connection-walk formalism that links self-attention to classical geometric operators, along with a set of operator-level tools for analyzing transformer models from a geometric perspective.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.