acceptodds
Under review as a conference paper at ICLR 2027

Copying Under Repetition: From Prefix Matching to Position Hashing

Abstract

Frontier LLMs that copy random strings accurately can overrun or truncate long blocks of repeated characters. Repetition removes the local content cue that induction-style copying relies on, and what is left to resolve—which occurrence to continue from, and when to stop—is positional. Position can reach the output along two routes: through the attention weights a head places on value vectors it leaves unchanged, or through offsets written into those value vectors, which a later layer can read and retrieve with. We study this distinction in horizon-uniform copying, which requires one transformer to copy every string of length without parameters specialized to the realized length. Our value-side construction, VO-rotation with exact access to modular residues, copies every string of length for any fixed finite vocabulary and all sufficiently large , using two layers, prime-period heads, and -bit precision. The same retrieval rule selects the next source token and the end marker, handling arbitrary repetition and termination together. We also prove a width-independent saturation bound: in our finite-precision model, transformers whose only positional input is an eventually-constant additive attention bias fail on all-identical sources once their saturation threshold is . Across three training seeds, learned VO models copy beyond the training limit, with an accuracy transition near the product of their supplied periods, consistent with learned use of the modular representation. In contrast, standard positional encoding baselines lose accuracy under repetition near the training limit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.