acceptodds
Under review as a conference paper at ICLR 2027

Relational Composition as Parallel Transport: What Per-Entity Frames Force in Attention

Abstract

A dot product means something only when both vectors share a coordinate system. Content words have one – the corpus fixes what every direction means – but a fresh entity, a name in a story or a variable in a program, does not: rotating its coordinates alone changes nothing. We take this per-entity frame freedom as a symmetry of the representation and ask what it forces in attention. Every invariant score depends on two entities only through their norms and the transported inner product , and every equivariant message lies in the span of and . A transport computed from the endpoints' own channels has rank at most the number of channels, never a group element from one vector each, so relations must be supplied; the standard transformer is the flat sector of one shared frame. Representation theory then attaches an integer to each axis of a mechanism, and the readout decides which integer. A readout with convex decision cells composes a finite group exactly only at widths of at least its smallest faithful real representation, and at that width every exact solution is a faithful representation with a finite latent group – as a conclusion, not a hypothesis. An even readout can decode a cover of the group instead, at a smaller width, and the gap is unbounded. Below its integer a mechanism is capped by the part of the group it cannot see – the commutator subgroup for a commutative mechanism, the derived series for a stack of commutative layers, the lower central series for a signature readout – all computable from the group before training. Empirically, every fitting seed at the minimal width generalizes exactly to length ( of , against of at twice the width), with certified horizons read from the weights. On , a -element group whose faithful width is , gradient descent at width learned the -state double cover under a quadratic readout and under sign-blind attention ( of seeds, half declared in advance; monotone attention of ), and on a frozen cover the cross-entropy-optimal affine head is provably uniform while a quadratic head is exact. On parsed CLUTRR, transported composition beats a matched transformer by points over four to ten hops (trained on two and three), and forcing isometries reverses that by – no isometry represents an idempotent rule such as sibling of sibling, as Krohn–Rhodes predicts – while on StepGame's commutative algebra isometries help instead. Static embeddings gave each word one vector; contextual models gave each occurrence its own; per-entity frames make each relation an operator, the unit in which relations compose. This points toward attention that produces its operators from context rather than retrieving them from a library, and the carrier theorem writes its specification in advance: not from the two entities, but from a carrier the context supplies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.