acceptodds
Under review as a conference paper at ICLR 2027

Disentangling Tokenizer Confounds in Bilingual Transfer Estimates

Abstract

Training a multilingual language model requires choosing how much data to use from each language, and a language trained alongside a target language can speed up or slow down its learning. Longpre et al. (2026) quantify this with the Bilingual Transfer Score (BTS): the number of tokens a 50/50 bilingual run needs, relative to a monolingual baseline, to reach the same target-language loss. Using a large set of empirical language modeling runs, they provide pairwise BTS estimates across many language pairs as a reusable resource for selecting data mixtures. Notably, all scores were estimated under a single multilingual tokenizer. It has been observed that intrinsic tokenizer metrics differ widely across languages (Petrov et al., 2023; Ahia et al., 2023), which raises the question: to what extent does a pair's score reflect the language pair vs. the tokenizer used to estimate it? Holding the model and training data fixed, we recompute BTS under tokenizers that differ in which of the two languages they are trained on, and analyze how the score changes with tokenizer training data composition, per-language fragmentation, and script. We confirm that tokenizer training data composition and language fragmentation are non-negligible factors systematically associated with the estimates. We then define the Tokenizer-adjusted BTS, which explicitly controls for tokenizer-induced variation. We re-estimate BTS for several language pairs using this method and show that it differs significantly from the original estimation method.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.