Tokenization Invariance Is Learned Late and Broken at the Head
Abstract
Splitting the input of a subword language model into single characters raises its bits per character five- to thirteen-fold, which is commonly taken to mean that the model is tied to the tokenization it was trained on. We argue that the failure is largely confined to the output head, and that what the trunk has learned survives re-tokenization on a schedule and with a specificity that current accounts do not predict. Probing fifteen pretraining checkpoints of OLMo-2-7B with frozen linear probes and matched random-direction control probes, we find that the model solves a word-order task within its first 9 billion tokens, but that the probe direction becomes reliably readable on character-split input, after a per-layer mean correction, only at about 630 billion tokens, roughly seventy times later. This invariance is specific to the task direction, increases with scale within OLMo-2 (one anomalous final checkpoint is flagged), and has not appeared by the end of two shorter Pythia runs. In released models it is graded in the fraction of tokens split and specific to the construct: no semantic-content probe in four models transfers more selectively than random directions in a way that survives a change of pooling rule, so re-tokenization spares positional structure in strong trunks rather than content. On the generation side, every model continues character-split prefixes at 1.3 to 1.8 times its canonical bits per character but emits character-split text at 5 to 13 times, and restricting the output distribution to character tokens removes at most a quarter of that collapse. Retraining the output head, 6.5% of the parameters, over the frozen trunk brings Llama-3.1-8B to within 0.63 bits per character of canonical; a capacity-matched adaptation of the trunk, run as a registered control, repairs less and does not preserve canonical reading. The part of the head repair attributable to the trunk tracks probe-measured invariance across eight models (exact one-sided Spearman , , registered in advance). Twenty predictions were registered before their data were collected; twelve failed, and every outcome is reported.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.