acceptodds
Under review as a conference paper at ICLR 2027

CSM-VL: Conditional Subspace Matryoshka for Vision-Language Embeddings

Abstract

Matryoshka Representation Learning (MRL) enables a single encoder to support multiple dimensionality budgets through nested embedding prefixes. However, conventional MRL uses the same dimensional ordering for every input, limiting its ability to adapt to varying information needs in vision-language embeddings. We introduce CSM-VL (Conditional Subspace Matryoshka for Vision-Language Embeddings), which replaces the fixed hierarchy with an input-dependent ordering of learned functional subspaces while preserving nested representations. CSM-VL partitions the embedding space into functional dimension groups and uses a conditional router to determine their order for each multimodal input. We further introduce Utility-Guided Group Allocation, which measures each group's contribution to the multimodal contrastive objective, and Conditional Multimodal Interaction, which encourages subspaces to capture informative cross-modal interactions. This allows different inputs to prioritize different subspaces under the same dimensional budget. Experiments across multimodal embeddings benchmarks and multiple vision-language backbones show that CSM-VL consistently improves the accuracy-dimensionality trade-off, particularly under constrained embedding budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.