acceptodds
Under review as a conference paper at ICLR 2027

OmniRouter: Routing, Not Scale A Multimodal Retrieval Encoder

Abstract

Multimodal retrieval requires a single space in which text, images, ambient audio and speech can be compared, and the dominant way to obtain one is to train a unified tower across all modalities. We ask whether such a space can instead be constructed from specialist encoders that were never trained together, and what the construction gains and loses. Two obstacles make this hard: the geometries of independently pretrained encoders are not compatible, so choosing any one of them as the shared space discards what the others capture, and practical systems truncate embeddings, which removes an appended second geometry entirely. OMNIROUTER addresses both. A closed-form ridge mapping between two text spaces defines a residual that holds what a vision-aligned encoder knows and a text encoder does not, and the two geometries are interleaved coordinate-wise so that every truncated prefix retains both. Because each modality reaches the space through its own projector, similarity decomposes per block and a change to one route leaves the scores of directions that do not involve it exactly unchanged. Both geometries contribute on every cross-modal direction, by up to 0.149 NDCG@10 over the better single one, and interleaving beats concatenation from 256 to 1024 dimensions. The system covers four input families with 23.5M trained parameters, 1.34% of a 1.76B model, and exceeds an 8.9B omni-modal embedder on all five core directions. Explicit routing also makes failures diagnosable: a speech family that looked like a capability limit was a missing dispatch rule, and one routing change raises its mean NDCG@10 from 0.0109 to 0.7657 with no retraining. Adding audio to a text-and-image system costs 150 seconds and changes no existing embedding, whereas adapting a jointly trained tower to the same data moves its text and image embeddings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.