acceptodds
Under review as a conference paper at ICLR 2027

DARE: Dual-Anchored Recurrent Reasoning for Universal Multimodal Embeddings

Abstract

Reasoning-enhanced universal multimodal embeddings have achieved strong performance by incorporating chain-of-thought (CoT) reasoning into representation formation. However, explicit CoT introduces substantial autoregressive decoding latency, while existing latent alternatives largely inherit token-wise rollout and provide limited constraints on how intermediate states should evolve and interact with visual evidence. This raises a fundamental challenge: how to efficiently internalize multimodal reasoning while maintaining semantic structure and faithful visual alignment? To this end, we propose DARE (Dual-Anchored Recurrent reasoning multimodal Embeddings). DARE performs block-wise recurrent reasoning in a fixed-width latent workspace, updating all latent tokens in parallel within each step and progressively deepening the reasoning trajectory across recurrent iterations, regulated by a dual-anchored mechanism during training: (i) a semantic trajectory anchor that associates recurrent states with ordered reasoning stages to prevent semantic drift; and (ii) a recurrent visual evidence anchor that preserves input-conditioned perceptual evidence across successive refinements through prototype level structural alignment and instance level contrastive matching. Both anchors are used only during training and introduce no auxiliary modules or inputs at inference time. Extensive experiments on 78 MMEB-V2 subsets spanning image, video, and visual-document tasks demonstrate that DARE outperforms previous reasoning-enhanced embedding models, while achieving 45 and 1.5 speedups over prior explicit-CoT and latent-reasoning approaches.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.