How Far Do Simple Transformations Translate Across Heterogeneous Text and Image Embedding Models?
Abstract
We investigate whether simple transformations can translate representations across heterogeneous text and image embedding models. Understanding how independently trained models organize semantic information is an enabler for AI-to-AI latent communication without decoding into human-readable text, and for reusing components across models. A recurring hypothesis in the literature is that independently trained models converge to a shared geometry over semantic content, so that a simple, low-supervision transformation should relate their internal representations; we test this hypothesis of latent universality in a realistic setting beyond simplified benchmarks. We study 14 text embedding models and 4 image/multimodal models that differ in architecture, pooling strategy, training objective, and dimension, and evaluate compatibility with linear CKA, downstream transfer, representation fidelity, and -NN retrieval, on full evaluation sets and three anchor protocols. Simple translators recover meaningful shared structure and support transfer for some compatible pairs, but fail sharply for others; same-design models cluster tightly while cross-design pairs degrade in structured ways. Compatibility depends jointly on architecture, training objective, readout/normalization, and data distribution. Overall, the results show that heterogeneous embedding spaces are not universally related by simple mappings as often suggested in the literature.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.