Back to the Roots? Rethinking Multilingual LLMs through Foundational Learner Vocabulary
Abstract
Multilingual large language models increasingly serve users across languages, yet their tokenizers fragment the lexical core: the words learners encounter first and rely on most. We introduce a new perspective that connects LLM vocabulary to language learner vocabulary, operationalized through LexCore, a 105-language foundational learner-vocabulary lexicon covering approximately 86.2% of the world's non-English-speaking population. Using LexCore, we study multilingual LLMs from diagnosis to intervention. Across 53 languages and 10 tokenizers, we show that contemporary tokenizers fragment most foundational learner vocabulary, with systematic disparities across resource levels, scripts, and language families. We further show that this fragmentation has a representation-level association: in bilingual retrieval, lexical grounding degrades as target words require more subword pieces, especially for lower-resource languages, while SUM emerges as the most stable composition operator. Building on these findings, we introduce LIFT, a learner-grounded, tokenizer-native vocabulary-injection method that reuses existing tokenizers, selects fragmented learner words under a fairness-aware maximin budget, and provides a worst-language relative-burden guarantee. LIFT improves several fairness-sensitive tokenization, fragmentation, and adoption metrics on FLORES+ and FineWeb2 without retraining tokenizers from scratch. Finally, in a controlled Gemma-3 4B continued-pretraining study, sum-initialized LIFT-3K achieves the best observed quality–parity–efficiency tradeoff among the controlled variants across multilingual QA, bidirectional translation, Information Parity, and decoding efficiency. These results show that foundational learner vocabulary is both a compact diagnostic lens for multilingual tokenizer inequity and a practical basis for targeted, budget-controlled adaptation of existing multilingual LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.