JoTS: Joint Token- and Segment-Level Prediction for Self-Supervised Multimodal Retrieval
Abstract
Multimodal dense retrieval learns representations for semantic matching across textual and visual content. However, existing multimodal retrievers largely rely on manually annotated or synthetically generated relevance pairs, overlooking the supervision naturally present in naturally occurring multimodal documents, where text and images form coherent sequences. This natural document order can be regarded as annotation-free supervision. We introduce JoTS (Joint Token- and Segment-Level Prediction), a self-supervised framework that jointly optimizes retrieval-conditioned next-token prediction (NTP) and segment-continuation prediction (SCP). NTP uses Retriever-weighted context to guide textual token prediction, while SCP identifies the succeeding text or image segment from a dynamic candidate set. These parallel objectives capture fine-grained linguistic dependencies and higher-level multimodal continuity. We pretrain JoTS on OBELICS using vision-language backbones of different scales. On general-purpose MMEB and reasoning-intensive MM-BRIGHT, JoTS consistently improves its base backbones and outperforms representative supervised and contrastively trained retrievers, demonstrating the effectiveness of predictive self-supervision for multimodal retrieval.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.