acceptodds
Under review as a conference paper at ICLR 2027

Separating operator geometry from learned routing in a gated, slot-routed associative memory

Abstract

Associative-memory and linear-attention architectures are often diagnosed through a functional of the binding operator: key orthogonality, the spectrum of AᵀA, a coherence bound on the feature map. We ask what training actually changes in a gated, slot-routed associative memory, in which each stored write reaches the logit through the product of a routing factor and an operator factor. Our main instrument is withheld learning: one factor is frozen at its initialisation, the rest is trained, and the task's loss is measured with every slot intact. (i) Routing. A frozen random hard hash costs 0.1 to 1.4 accuracy points (three seeds), where a frozen diffuse router costs 37 to 55: a concentrated, balanced assignment supplies most of what routing contributes. (ii) Writes. Frozen write magnitudes cost 13 to 28 points (about 5 once token identity no longer marks a write's type; two loads, one seed); frozen codebook directions cost 2 to 6. (iii) Operator geometry. Reading the trained model agrees, once the reading is a function of the model. The customary orthogonality deviation is not: it moves by up to 255× along a gauge orbit, a change of basis that leaves every logit unchanged. On a gauge-invariant contrast the operator's cross-key arrangement sits at its random-binding floor at every slot count and binding dimension we train, and no search that holds accuracy within two points lifts it past 1.03×. Within routing, concentration (how few slots a write lands on) is the lever, and arrangement (which writes collide) is a small, measurable channel: its mean square slot overlap sits a few per cent below permutations matched on concentration, and a fraction of one per cent below permutations matched on concentration and load. A relation joins the two readings. Under a Gaussian-crosstalk model whose variance assumption is checked, descriptors read off the forward pass (routing weights, write norms, same-key overlaps) predict retrieval accuracy to within 0.08 across 124 models with nothing fitted, so the reading is a measurement other scales can repeat rather than an extrapolation to them. The analysis stops where the memory forgets: on DeltaNet the static coefficient of a stored item does not exist, and the transition that removes it has a Lyapunov spectrum we measure. We do not offer a capacity law.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.