acceptodds
Under review as a conference paper at ICLR 2027

Improving Multilingual Semantic Alignment with Multi-Positive Contrastive Learning for Information Retrieval

Abstract

Multilingual embedding models do not always retrieve all semantically equivalent relevant documents within the top-ranked results. For a given query, relevant documents in some languages are ranked substantially lower than parallel documents written in other languages, and retrieval performance also varies across query languages. In this paper, we introduce Multilingual Semantic Alignment (MuSA), a training method that aims to enhance semantic similarity among parallel document embeddings while jointly optimizing retrieval over multilingual positives associated with each query. Beyond query–document relevance, MuSA directly aligns the embeddings of multilingual parallel documents through a dimension-wise cross-correlation alignment objective. Experiments with 13 multilingual embedding models show that MuSA improves multilingual retrieval performance in settings with multiple references and reduces performance disparities across query languages observed in the baselines. Across the evaluated settings, MuSA reduces the depth required to retrieve all relevant documents by up to 84.62% relative to the corresponding baselines. Analysis shows that improvements in multilingual retrieval are associated with better alignment of parallel document embeddings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.