Mind the Gaps: Learning Positions That Generalise
Abstract
Transformers can solve algorithmic tasks nearly perfectly within the domain; however, they typically fail on out-of-distribution evaluation, specifically on longer input sequences than those seen during training. This lack of generalisation has been partially addressed in the literature by assigning token positions according to their computational roles. Removing this guidance is challenging: the decoder can adapt to learned positional assignments that support accurate predictions within the training range but fail to generalise to longer sequences. Here, we show that models can learn positional assignments that support length generalisation from next-token prediction alone. To make these assignments easier to identify, we train across randomly gapped coordinate systems. These systems preserve exact positional matches, whereas mismatches induce varying displacements that are harder for the decoder to compensate for consistently. Building on this principle, we introduce Reference-relative RoPE (ReRoPE), which learns a mixture of reference-relative RoPE transformations. For each token, ReRoPE scores the nearest visible occurrences of every token type as candidate references, with each candidate defining a transformation in the mixture. Across addition, copy, reverse, and parity tasks, ReRoPE generalises to lengths several times the training maximum, matching supervised position coupling. Further experiments on addition show that ReRoPE retains this performance in deeper models and exceeds position coupling in terms of the maximum length to which the model generalises.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.