Retokenization reveals a gap between semantics and behaviror
Abstract
Language models are trained under a deterministic canonical tokenization, yet every byte string admits exponentially many non-canonical token sequences that decode to it. Although such sequences never occur during training, models exhibit partial invariance to them, and in some settings, such as character-level word games, non-canonical inputs even improve accuracy. This invariance is nevertheless fragile: byte-identical prompts can induce divergent behavior, and adversarial tokenizations were found to circumvent safety alignment. We give a mechanistic account of this fragility and introduce a consistency-training objective that enforces agreement across retokenizations of the prompt while keeping the output space canonically tokenized. The resulting models are more robust to adversarial tokenization and exhibit reduced residual-stream sensitivity to segmentation differences. We further show that retokenization cannot serve as a substitute for prompt diversity: although retokenizations of a prompt may share no common tokens, training on them fails to reproduce the generalization gains of a comparable increase in the number of distinct byte strings. Generalization thus appears to be governed by the diversity of latent representations rather than of surface token sequences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.