CSM-VL: Conditional Subspace Matryoshka for Vision-Language Embeddings
Abstract
Matryoshka Representation Learning (MRL) enables a single encoder to support multiple dimensionality budgets through nested embedding prefixes. However, conventional MRL uses the same dimensional ordering for every input, limiting its ability to adapt to varying information needs in vision-language embeddings. We introduce CSM-VL (Conditional Subspace Matryoshka for Vision-Language Embeddings), which replaces the fixed hierarchy with an input-dependent ordering of learned functional subspaces while preserving nested representations. CSM-VL partitions the embedding space into functional dimension groups and uses a conditional router to determine their order for each multimodal input. We further introduce Utility-Guided Group Allocation, which measures each group's contribution to the multimodal contrastive objective, and Conditional Multimodal Interaction, which encourages subspaces to capture informative cross-modal interactions. This allows different inputs to prioritize different subspaces under the same dimensional budget. Experiments across multimodal embeddings benchmarks and multiple vision-language backbones show that CSM-VL consistently improves the accuracy-dimensionality trade-off, particularly under constrained embedding budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.