Self-Bootstrapped Multimodal Embeddings for Robust Text-to-Multimodal Retrieval
Abstract
Text-to-multimodal retrieval aims to retrieve documents containing both images and text in response to textual queries. Existing multimodal embedding models typically rely on large-scale heterogeneous training corpora, external teachers, or auxiliary data-generation models, while task-specific text-to-multimodal training data remain scarce. We present Self-Bootstrapped Multimodal Embeddings, a framework that exploits the backbone MLLM’s own generation, reasoning, and retrieval capabilities to improve multimodal representations without external teachers or auxiliary annotation models during retrieval training. First, the backbone generates detailed descriptions for document images, self-annotating existing text-to-image corpora and converting them into text-enriched multimodal retrieval data. Second, reasoning instruction conditioning during contrastive training elicits and internalizes the model’s native reasoning-oriented computation into single-vector embeddings, requiring neither explicit reasoning traces nor additional reasoning tokens at inference. Finally, the trained retriever uses its own retrieval capability to identify highly ranked non-relevant documents and adopts them as self-mined hard negatives for targeted representation refinement. The resulting framework improves the average performance of both Qwen3-VL-Instruct and VLM2Vec-V2 on text-to-multimodal and text-to-video retrieval. With only 300K training pairs, our VLM2Vec-V2-based model achieves an average Recall@1 of 81.3% across six text-to-multimodal benchmarks—the best among models fine-tuned with no more than 300K task-specific pairs—and reaches 94.1% on EDIS. We further introduce EDIS-Noisy and WebQA-Noisy to simulate realistic retrieval scenarios. Without noise-specific training, our model achieves average Recall@1 scores of 91.4% and 80.4%, respectively, outperforming both Qwen3-VL-Embedding and WeMM-Embedding. These results demonstrate that self-bootstrapping the backbone’s existing capabilities provides an effective approach to data-efficient and robust multimodal retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.