acceptodds
Under review as a conference paper at ICLR 2027

Cross-Lingual Data Scaling for Large Language Models

Abstract

Many target languages have limited pretraining text even as English corpora continue to grow. We study a controlled cross-lingual scaling question: when the available target-language corpus is fixed, how does target-language performance change as English data increases? In English–Chinese experiments, proportional mixing dilutes Chinese exposure and can worsen Chinese perplexity. Under staged training, a single cosine schedule couples the Chinese-stage learning rate to the length of English training; holding the Chinese-stage schedule fixed instead produces positive trends over the evaluated scales. Token-category ablations are consistent with shared lexical content contributing to this transfer, although they do not establish a general mechanism. Repeating the fixed target-language corpus preserves its effective training share, with diminishing returns at high repetition. We instantiate these interventions in ScaleX and ScaleX-DR and evaluate them across English data scales. In multilingual experiments that scale English while holding Chinese, Turkish, Hungarian, and Bengali corpora fixed, ScaleX-DR attains lower perplexity and higher downstream accuracy at the largest than at the smallest English scale for all four target languages. A separate 2.5B-parameter English–Chinese comparison at 500B and 1T tokens improves both languages. These results characterize how schedule and target-language exposure shape cross-lingual scaling in the evaluated settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.