acceptodds
Under review as a conference paper at ICLR 2027

Measuring the Language Share of Multilingual Representations

Abstract

Multilingual language models can represent the same content in different languages, but the balance between language and sentence differences inside these models remains difficult to measure. We introduce Language Factor Share (LFS), which uses translations of the same sentences to measure average language differences relative to sentence differences at each layer, without choosing a reference language. Across 33 open base models, language share is lowest at an intermediate layer and higher at the input and output. This pattern repeats on a second collection of translations for 17 models, and the decrease varies widely among similarly sized models. Lower language share is also associated with higher average accuracy on two multilingual benchmarks. In a preregistered study of 19 models, lower language share goes with greater benefit from relevant context in another language when predicting English text, after accounting for a proxy for language data availability, language family, script, and tokenization. Reporting the two parts the share is built from shows which one changed, or whether both did. Our training experiments did not establish that lowering the share improves performance. LFS therefore helps researchers examine how multilingual representations are organized and interpret their changes alongside direct tests of model performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.