AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval
Abstract
Multi-vector representations have emerged as an effective paradigm for multimodal retrieval, where each sample is encoded by multiple complementary embeddings to capture fine-grained cross-modal information. However, existing approaches typically use a fixed representation capacity, assigning the same number of vectors to every sample regardless of its content or retrieval difficulty. This fixed-capacity design overlooks the fact that different samples may require different amounts of representation capacity for effective retrieval. In this work, we introduce Sample-Adaptive Multi-Vector Representation (SAMVR), a new problem setting for multimodal retrieval that studies how multi-vector representation capacity can be allocated at the sample level. Under SAMVR, each sample is represented by a content-adaptive embedding set (CAES), whose size is determined according to the sample-specific retrieval utility of additional representation vectors. To instantiate SAMVR, we propose AdaptiveEmbed, a unified framework for learning sample-adaptive multi-vector representations. AdaptiveEmbed learns structured multi-vector representations through Multi-Group Contrastive Learning (MGCL) with the symmetric set-to-set similarity (SetSim), and further employs Utility Policy Optimization (UPO) to determine sample-specific representation capacity via Marginal Utility Allocation (MUA). Experiments across multimodal retrieval benchmarks involving image, text, video, and audio show that sample-adaptive capacity allocation achieves overall better retrieval performance than fixed-capacity multi-vector representations, validating SAMVR as a viable formulation for multimodal retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.