acceptodds
Under review as a conference paper at ICLR 2027

An Exact Advection-Diffusion-Reaction Representation of Transformers

Abstract

Repeated Transformer architecture invites a dynamical-systems view, but leaves open how computation is organized and changes through depth. Building on recent work connecting self-attention to a Laplacian-like difference operator, we systematically analyze how attention, normalization, residual connections, and nonlinear feed-forward transformations compose. We obtain an exact advection–diffusion–reaction (ADR) representation of sequential pre-norm Transformer blocks, retaining the feed-forward response at the attention-updated state. This representation provides a common geometric framework for quantifying component actions, deriving constraints on causal modifications, and examining task behavior through controlled interventions. In the evaluated pretrained decoders, advection generally grows relative to the complete update while reaction's relative magnitude declines, despite absolute reaction growth in most checkpoints. On common inputs, neighboring complete updates are more aligned than a fixed distant-block control, while their transport and reaction actions remain directionally distinct. Separating the identity baseline reveals that apparent diffusion persistence can coexist with changing symmetric transport. An attention-approximation intervention further shows that an opposing feed-forward response can amplify the block-output error. By preserving the complete nonlinear computation, exact ADR provides a quantitative basis for relating component geometry to depthwise behavior and the effects of modifying a block. This connection supplies constraints for architecture design and evaluation criteria for inference approximation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.