MATE: Matching-Aware Latent Reasoning for Universal Multimodal Embedding
Abstract
Universal multimodal embedding (UME) enables scalable retrieval by encoding queries and candidates independently in a shared vector space. Complex retrieval can nevertheless benefit from integrating evidence within each input and directly comparing queries with semantically similar candidates. Explicit reasoning and multi-step latent rollout increase encoding cost, while pairwise matching adds per-candidate computation at inference. We propose MATE, a one-step bi-encoder that learns from both forms of computation. Latent Reasoning Distillation (LRD) uses answer-guided weighting to summarize a teacher's multi-step latent trajectory and progressively distills this target into a single latent step. Latent-Aware Matching Distillation (LAMD) trains a matcher on paired latent states and mean-pooled embeddings, then selectively transfers its listwise preferences to the bi-encoder. Matching distillation focuses on queries with a sufficient matcher margin but an insufficient bi-encoder margin. At inference, MATE encodes queries and candidates independently and ranks candidates by normalized vector similarity. On the 78-task MMEB-V2 benchmark, the 2B MATE model achieves an overall score of 65.4 and delivers 62 the throughput of the explicit-reasoning model UME-R1 under matched inference settings. Code will be publicly available upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.