MM-DiffEmbed: When Multimodal Diffusion LLMs Meet Multimodal Document Retrieval
Abstract
Multimodal document retrieval (MDR) requires representations that preserve local evidence while interpreting it in context. *Most recent multi-vector retrievers derive these representations from autoregressive multimodal language models, leaving the role of diffusion-pretrained contextualization in this setting unclear*. We introduce **MM-DiffEmbed**, a framework that adapts multimodal diffusion language models (MDLMs) into single-pass, multi-vector retrievers, instantiated as ColLaViDa and ColSDAR-VL. The framework retains contextualized token embeddings rather than pooling an entire page into one vector, and supports prompt-denoising and query-vector decorrelation during retrieval adaptation. We evaluate the models on ViDoRe-v1, ViDoRe-v2, and MRMR, a reasoning-intensive MDR benchmark. On MRMR under a shared evaluation protocol, ColLaViDa and ColSDAR-VL exceed their strongest respective LLaVA- and Qwen-family baselines by 3.79 and 4.76 points. These gains coexist with a task-dependent trade-off: ColLaViDa slightly exceeds ColLLaVA-Llama3 on ViDoRe-v2, whereas ColSDAR-VL remains below ColQwen on both ViDoRe versions. Overall, the results establish a concrete testbed for studying diffusion-pretrained multi-vector encoders in future MDR research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.