TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
Abstract
Universal multimodal retrieval must map both explicit matches and unresolved compositional intents into a shared embedding space. Intermediate reasoning can clarify a query's target semantics, but generating it for every input adds latency and may weaken retrieval. We introduce TRACE, a framework that learns retrieval representations through selective query-side reasoning. Its first autoregressive prediction selects immediate embedding or a retrieval-oriented reasoning sequence; both routes produce a vector from the hidden state predicting the same embedding marker. Joint generation and contrastive training couple intent resolution to representation learning, while directly encoded candidates remain indexable offline. We construct M-BEIR-CoT with 518K direct examples and 575K filtered reasoning examples to supervise both behaviors. TRACE achieves a 58.8 average score on M-BEIR, improving a same-backbone LamRA-Ret reproduction by 2.1 points, and improves over LamRA-Ret in 11 of 13 distribution-shifted settings. On separate MSCOCO and CIRR controls, it attains higher recall than independently trained direct-only and always-reasoning policies, with approximately the query throughput of always reasoning. CIRR interventions and ablations further support reasoning-enabled inference, joint representation learning, and query-only generation. Together, these results demonstrate selective reasoning within a shared, index-compatible multimodal retriever.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.