acceptodds
Under review as a conference paper at ICLR 2027

Read in English, Answer in Hindi? Exploring Data Compositions for Cross-Lingual Knowledge Transfer

Abstract

Recent evidence suggests that multilingual language models (LMs) can inexplicably transfer knowledge acquired in one language to other languages without explicit incentives for such behavior. Factors that govern such transfer aren’t well understood, with past studies disagreeing on what constitutes knowledge transfer and what factors are important. Our work addresses this gap, and studies the role of patterns in multilingual pre-training data in unlocking cross-lingual knowledge transfer. We devise a controlled setup wherein we train language models from scratch on a synthetic corpus comprising biographies across multiple languages. Our experiments highlight that introducing parallel data improves transfer asymmetrically for dissimilar languages. Additionally, how parallel data is introduced also matters: smaller amounts of parallel data across many documents is more beneficial than higher parallelism across fewer documents. In addition to parallel data, the presence of code mixing across documents or complementary information across languages also improves transfer. We find that these findings also hold in real-world settings, where we continually pre-train off-the-shelf models on new knowledge, albeit transfer is moderate. The trends generalize similarly for a two-hop question-answering task, which requires a mix of cross lingual recall and reasoning. Overall, our findings offer data-curation strategies for pre-training that enable knowledge transfer in multilingual language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.