acceptodds
Under review as a conference paper at ICLR 2027

ETBench: Evaluating Embedding Translation for Mixed-Index Retrieval

Abstract

Mixed-index retrieval searches documents encoded by different embedding models, e.g., a legacy model and its replacement. Embedding translation enables this by mapping stored source embeddings into the target space, where queries are also encoded. Translation often uses embedding alignment, as in bilingual lexicon induction, fitting a map to anchor pairs, i.e., the same documents encoded by both models, and applying it to the remaining source-only vectors. We show that better alignment does not necessarily yield better mixed-index retrieval, and prove a sufficient and necessary condition under which alignment maps can preserve the target model's top-k set. We further develop ETbench, a benchmark evaluating alignment methods for embedding translation across 24 directed encoder pairs and four corpora. It tests whether translation makes source-only documents searchable alongside native target embeddings in a mixed index. Our results show that high alignment accuracy can coexist with poor retrieval of translated documents, and common retrieval measures can be insensitive to this.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.