acceptodds
Under review as a conference paper at ICLR 2027

Parity Echoes in a Transformer: Predictable Far Length Re-emergence After Single Length Training

Abstract

We study how a Transformer trained to solve parity at a single sequence length behaves when evaluated on other lengths. Surprisingly, failure is not monotonic: the model can lose measurable parity signal at nearby unseen lengths and partially recover it much farther away in a regular pattern. We show that this behavior is closely related to how the model represents its inputs. Successful models organize sequences mainly by the number of ones they contain, with different counts arranged along an approximately one-dimensional structure. This representation makes the model’s behavior at unseen lengths predictable: we derive where parity signal should reappear and how its strength changes with sequence length, and find close agreement with trained models. We also study how the difficulty of learning parity grows with sequence length. Empirically, the number of optimizer steps required to learn the task scales close to , and a simplified spectral model yields the same scaling. These results connect two aspects of parity learning: how a trained model behaves away from its training length, and how the difficulty of learning the task grows as that length increases.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.