acceptodds
Under review as a conference paper at ICLR 2027

MAPP: Cross-Lingual Transfer through Representation Alignment with Performance-Gap Probing

Abstract

Cross-lingual representation alignment — training a Large Language Model (LLM) to encode semantically equivalent inputs similarly across languages — is increasingly used to improve LLM performance on low-resource languages, motivated by the observation that a language's performance gap to English correlates with the distance between their hidden states, particularly at middle layers. State-of-the-art alignment methods optimize distance functions or contrastive losses selected a priori, all of which are minimized in the limit of complete alignment. We argue that these objectives are only proxies for optimal cross-lingual transfer learning by showing that complete alignment both eliminates language identity and yields sub-optimal performance gains. By training oracle models and examining their multilingual representation geometry, we establish the performance gap and the language gap as separable components of cross-lingual representation difference. We therefore propose MAPP (Multilingual Alignment through Performance-gap Probing), which trains probes on middle-layer hidden representations to predict cross-lingual performance gaps and uses them as reward models to align the LLM. Across three tasks, MAPP yields larger performance gains than state-of-the-art alignment signals while preserving output language identity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.