acceptodds
Under review as a conference paper at ICLR 2027

From Explicit CoT to Relevance Guidance: Discriminative Latent Reasoning for Universal Multimodal Embedding

Abstract

Reasoning has shown promise for improving universal multimodal embeddings, through both explicit chain-of-thought (CoT) and CoT-supervised latent reasoning. However, explicit CoT may not reliably convey the discriminative information that embeddings rely on to preserve distinctions among inputs. We argue that such information can be acquired and internalized by reasoning without necessarily relying on explicit CoT. To realize this view, we introduce Discriminative Latent Reasoning for Universal Multimodal Embedding (DiLRE), which establishes a discriminative foundation for latent reasoning through contrastive supervision and internalizes complementary information through relevance guidance. We train a Discriminative Guidance Model to jointly assess query–target relevance, providing guidance that complements embedding similarity. At test time, this guidance is combined with the embedding's current retrieval distribution to form discriminative targets for latent refinement. An iterative loop of retrieval and latent refinement then updates the latent reasoning tokens to align the resulting embedding's retrieval distribution with these targets. Extensive experiments and ablation studies demonstrate the effectiveness of DiLRE across diverse multimodal embedding tasks and model scales.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.