acceptodds
Under review as a conference paper at ICLR 2027

Semi-Continuous Representation Learning: Mixed-Width Contrastive Training for Calibrated Matryoshka Embeddings

Abstract

Matryoshka Representation Learning (MRL) has become a widely adopted approach for preserving semantic structure across embedding dimensionalities. Its premise is that each Matryoshka prefix is a coarser representation of the same input within one common embedding space. However, while all nested embeddings share the same space, their similarity scores are miscalibrated across prefix sizes. We demonstrate theoretically and experimentally that the miscalibration stems from MRL's loss formulation where each prefix is optimised under its own partition function. Therefore, we propose Semi-Continuous Representation Learning (SCRL), which replaces MRL's decomposed loss with a single contrastive objective over randomly sampled prefix sizes under a shared softmax. We prove that SCRL results in calibrated prefixes and show empirically that it ensures consistent similarity scores across nested dimensions, improving mixed-size retrieval performance by up to 8.8 nDCG@10 points on average across eight BEIR benchmarks. Furthermore, SCRL improves cross-size neighborhood consistency and increases the score distribution overlap between 64- and 768-dimensional embeddings from 71.7% to 84.9%. SCRL can be applied as a direct training replacement for MRL, as a post-training step for existing MRL embedding models, and in MRL-adaptor settings on top of any general-purpose embeddings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.