acceptodds
Under review as a conference paper at ICLR 2027

Transformers as effective fields: from symmetry breaking to architecture

Abstract

Can Transformer-like structures be derived from structural assumptions on a learnable field? We formulate an axiomatic effective-field theory that does not postulate attention in advance. For analytic matrix potentials with token-isometry invariance and feature-basis covariance, we classify the Taylor expansion through fourth order. The quadratic term generates token-wise linear feature mixing, while the quartic Hessian generates a cubic interaction with a finite paired linear-attention decomposition. This pairing reflects the trace-self-adjoint structure imposed by a scalar potential; relaxing the exchange tying opens a trace-anti-self-adjoint, locally non-potential sector that accommodates generic untied attention. A separate low-complexity assumption yields a small number of bottleneck heads. Row-softmax weights remain token-permutation equivariant and admit a row-wise entropy-regularized variational characterization, but softmax normalization does not in general restore potential compatibility for an untied update. We further quantify how a finite Transformer-like model tracks a fixed effective flow. Under uniform regularity, its trajectory error separates into an \(O(T^-1)\) depth-discretization term and width-dependent best field-approximation errors for attention and MLP. Controlled experiments diagnose these terms through frozen-field depth refinement and learned field surrogates. The field perspective therefore provides a low-order route from symmetry to architecture, with explicit limits on potential compatibility and approximation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.