acceptodds
Under review as a conference paper at ICLR 2027

Rotor Recurrences: What Compact Recurrent States Can Predict

Abstract

Compact recurrent states can preserve order-sensitive computations, but what a state can store need not be what a continuous decoder can recover. For pure rotor transitions associated with rotations in three or four dimensions, we prove a sharp prediction limit for any fixed finite number of heads. Under i.i.d. inputs from a fixed distribution whose target labels generate the group, no fixed continuous strictly positive decoder can asymptotically achieve expected prefix-averaged log loss below (\log 60) when predicting permutations of five objects. The bound is sharp: a parity-tracking recurrence approaches this floor with a decoder that assigns almost all probability mass uniformly within the correct parity class. The same rotor family nevertheless permits arbitrarily small loss when tracking even permutations alone. A general compact-group result characterizes the hidden ambiguity underlying these limits. We then study related affine rotor models with additive writes and decay, which lie outside the theorem. A single rotor layer learns the even-permutation task and generalizes to sequences four times the training length. Across implemented parameterizations of the same ideal rotation family, median sustained solve times among successful runs span roughly a factor of five. These experiments measure task learning without identifying the learned internal representation. Together, the results distinguish what recurrent dynamics can represent, what continuous decoding can recover, and what gradient training discovers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.