SEPARATING SCRIPT AND VOCABULARY EFFECTS IN MULTILINGUAL TOKENIZATION COST
Abstract
Cross-lingual token cost depends both on a language's writing system and on how its tokenizer encodes that writing system. We show that token-count ratios on parallel text factor exactly into these two components, and measure them across seven production tokenizers, fifteen languages, and five script families on FLORES. This decomposition distinguishes languages that are expensive because their writing system requires more characters from those with inefficient tokenization. Devanagari and Arabic-script languages require two to three times as many tokens as English despite similar writing-system cost, with most of the excess attributable to the tokenizer (convention-dependent for Devanagari: 99.9% under code points, 76.9% under grapheme clusters, though the cost ratio is not). We then measure, in simulation, how much of this tokenizer penalty can be recovered by replacing unused entries with high-impact fragments at fixed vocabulary size. At 5,000 slots, mean characters-per-token improvement reaches 58.8% for Devanagari and 51.1% for Arabic-script languages (per-language means over the seven tokenizers, evaluated on held-out devtest), with a maximum of 159.9%. CJK languages have comparable tokenizer penalties, but a segmentation-free generator with supply well in excess of the budget recovers almost nothing more. These results show that multilingual tokenization disparities should first be diagnosed by cause: substantial inefficiency can often be recovered by reallocating existing vocabulary capacity, but fixed-vocabulary interventions have clear limits.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.