Whose Tokenizer? When and How Much Tokenizer Choice Matters for Multilingual Language Models
Abstract
Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages. To this end, we train 123 language models across 54 tokenizers. Our main comparison spans 54 tokenizers; architecture, training corpus, training-token budget, and optimization are held fixed across the corresponding models, allowing us to isolate differences due to the tokenizer. We measure every outcome per language. Tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho=-0.54 over the 31 trained languages and -0.69 over the 28 written with word boundaries). Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data. Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages: both equal weighting and an allocation inverted against their language model training data share increase their BPB, particularly when language-model training repeats data. Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language. Over the 31 languages, we achieve 0.89 accuracy on pairwise rankings of trained models' mean BPB using just tokenizer intrinsic metrics, suggesting a practical strategy for screening tokenizer candidates before training language models. All resources will be released upon de-anonymization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.