Controlling representation sharing in multilingual transformers: a mechanistic study
Abstract
Understanding the mechanism through which transformers represent shared concepts across languages is a central problem in the study of multilingual large language models. Building on previous work, we find further evidence that multilingual LLMs encode the same concepts with similar representations across languages. To precisely understand shared representations in transformers, we analyze a single-layer transformer trained on a multilingual modular addition task, in which distinct vocabularies express the same underlying numerical concepts. We provide a mechanistic explanation of how these models implement arithmetic using shared Fourier representations across languages. We design a task-specific intervention on the training process of the multilingual modular addition model. This label-based intervention encourages representation sharing across languages and substantially improves cross-lingual generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.