Closing the Trajectory Gap: On-policy Learning For Multimodal Reasoning Embeddings
Abstract
Universal multimodal embeddings support a broad range of applications through a shared representation space. Although they build on multimodal large language models, most existing methods use these models primarily as encoders, leaving their generative capabilities underexplored. Recent reasoning-enhanced approaches generate intermediate reasoning before producing the final embedding and have shown promising performance on multimodal embedding benchmarks. However, training a unified reasoner–embedder often relies on task-specific chains of thought generated and filtered offline, while jointly optimizing language modeling and contrastive learning on these fixed trajectories. Beyond the cost of constructing this offline supervision, the model is trained on externally generated trajectories but conditions its inference-time embeddings on its own generations, creating a reasoning trajectory gap. Our controlled experiments suggest that this gap can undermine retrieval performance. We introduce OPR-Embed (On-policy Reasoning Embedding), a two-stage on-policy framework for multimodal reasoning embeddings. Without any pre-built CoT corpus, Stage I performs embedding-attentive on-policy distillation by scoring prefixes sampled from the current student with a frozen teacher and weighting token-level supervision using attention from the final embedding readout. Stage II uses verifier-guided group relative policy optimization to refine complete trajectories using retrieval feedback together with a candidate-match verifier. Experiments on MMEB-v2 demonstrate that OPR-Embed is competitive with state-of-the-art methods, achieving overall scores of 65.8 and 70.0 with 2B and 7B backbones, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.