OmniRouter: Routing, Not Scale A Multimodal Retrieval Encoder
Abstract
Multimodal retrieval requires a single space in which text, images, ambient audio and speech can be compared, and the dominant way to obtain one is to train a unified tower across all modalities. We ask whether such a space can instead be constructed from specialist encoders that were never trained together, and what the construction gains and loses. Two obstacles make this hard: the geometries of independently pretrained encoders are not compatible, so choosing any one of them as the shared space discards what the others capture, and practical systems truncate embeddings, which removes an appended second geometry entirely. OMNIROUTER addresses both. A closed-form ridge mapping between two text spaces defines a residual that holds what a vision-aligned encoder knows and a text encoder does not, and the two geometries are interleaved coordinate-wise so that every truncated prefix retains both. Because each modality reaches the space through its own projector, similarity decomposes per block and a change to one route leaves the scores of directions that do not involve it exactly unchanged. Both geometries contribute on every cross-modal direction, by up to 0.149 NDCG@10 over the better single one, and interleaving beats concatenation from 256 to 1024 dimensions. The system covers four input families with 23.5M trained parameters, 1.34% of a 1.76B model, and exceeds an 8.9B omni-modal embedder on all five core directions. Explicit routing also makes failures diagnosable: a speech family that looked like a capability limit was a missing dispatch rule, and one routing change raises its mean NDCG@10 from 0.0109 to 0.7657 with no retraining. Adding audio to a text-and-image system costs 150 seconds and changes no existing embedding, whereas adapting a jointly trained tower to the same data moves its text and image embeddings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.