acceptodds
Under review as a conference paper at ICLR 2027

Orthogonal Maps Are Sufficient to Align Multimodal Contrastive Models

Abstract

As models and data scale, independently trained networks often learn similar representations. Yet establishing similarity is weaker than identifying an explicit correspondence between representation spaces, especially for multimodal models, where consistency must hold both within modalities and across the learned coupling between modalities. We therefore ask whether two independently trained multimodal contrastive models, with encoders and , trained on different data distributions and with different architectures, exhibit a systematic geometric relationship between their embedding spaces. Across model families including CLIP, SigLIP, FLAVA and CLAP, we find that this relationship is well approximated by a single orthogonal transformation, i.e., there exists an orthogonal map , with , such that for paired images (audio) . Strikingly, the same map also aligns the text encoders, so that for text inputs . Theoretically, we prove that if two models agree on the multimodal kernel over a small finite anchor set, meaning , then their image (audio) and text representations must be related by a single orthogonal map. We validate this premise, showing strong qualitative and quantitative kernel agreement. This geometric alignment enables backward-compatible model upgrades without costly re-embedding, and has implications for the privacy of learned representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.