When Alignment Becomes an Interface: Deployment-Dependent Interoperability of Text Embeddings
Abstract
Search indexes can outlive the embedding models that created them, leaving documents represented by different encoders. Alignment makes these vectors comparable, but does a method that works well for one document encoder also work when several encoders' documents compete? We study seven frozen text encoders, comparing linear alignment with a shared nonlinear correction across four freshly encoded retrieval datasets and stored vectors for 2.68 million Natural Questions documents. Each document contributes only one encoder's vector. The same correction lowers retrieval quality when all documents use one encoder, yet raises it when documents use a balanced mixture of encoders. On a prospective duplicate-question test, the top-10 relevance score falls by 0.016 in the first setting and rises by 0.103 in the second, on a 0-1 scale. Exchanging score values while preserving each source's document order shows how better competition between sources can outweigh worse ordering within them. The prospective test supports both patterns after adjustment for multiple comparisons. A simpler alternative, standardizing each query's scores against reference texts, improves linear alignment on all four fresh datasets. However, it generally trails nonlinear correction in quality and takes more search time locally. These findings show why alignment should be evaluated on the index it will serve: preserving order within one source and comparing scores across sources are distinct requirements with different quality and cost tradeoffs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.