acceptodds
Under review as a conference paper at ICLR 2027

Retokenization reveals a gap between semantics and behaviror

Abstract

Language models are trained under a deterministic canonical tokenization, yet every byte string admits exponentially many non-canonical token sequences that decode to it. Although such sequences never occur during training, models exhibit partial invariance to them, and in some settings, such as character-level word games, non-canonical inputs even improve accuracy. This invariance is nevertheless fragile: byte-identical prompts can induce divergent behavior, and adversarial tokenizations were found to circumvent safety alignment. We give a mechanistic account of this fragility and introduce a consistency-training objective that enforces agreement across retokenizations of the prompt while keeping the output space canonically tokenized. The resulting models are more robust to adversarial tokenization and exhibit reduced residual-stream sensitivity to segmentation differences. We further show that retokenization cannot serve as a substitute for prompt diversity: although retokenizations of a prompt may share no common tokens, training on them fails to reproduce the generalization gains of a comparable increase in the number of distinct byte strings. Generalization thus appears to be governed by the diversity of latent representations rather than of surface token sequences.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.