acceptodds
Under review as a conference paper at ICLR 2027

String Atlas: Breaking the One-Tokenizer Assumption in Language Modeling

Abstract

A standard language model is trained with a single canonical tokenizer, so that tokenizer is effectively frozen into its weights. This structural constraint means that changing the tokenizer afterwards invalidates the learned embedding parameters connecting raw text to the Transformer trunk. Yet, tokenizer choice is far from a cosmetic preprocessing step; it dictates domain compression, effective context length, arithmetic capabilities, and multilingual fairness. While existing methods either commit to one universal vocabulary upfront or resort to expensive post-hoc vocabulary transfer, we introduce String Atlas, an architecture that natively supports multiple tokenizers from pretraining. Instead of equipping each tokenizer with private embedding tables, which fragments learning, String Atlas keys its lexical parameters directly to the emitted byte string. Tokenizers may disagree on boundaries, but they seamlessly share rows in a common atlas whenever they emit the same string. Across suites of tokenizer sets varying domains, granularity, and even algorithm, we find that String Atlas allows a language model to support several tokenizer interfaces simultaneously, nearly matching the performance of independently trained specialists while maintaining the training and inference costs of a single model. Together, these results break the one-tokenizer assumption that binds a model to its vocabulary and reframe tokenizers as interchangeable views over a single shared lexicon.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.