MORSE: Mixture of Representations from Sparsified Experts
Abstract
Recent work on quantization (e.g., TurboQuant) and dimensionality reduction (e.g., Matryoshka Representation Learning; MRL) demonstrates that text embeddings contain substantial representational redundancy. However, enforcing extreme sparsity (retaining only a small number of active dimensions) can reduce embedding generality and favor task-specific performance, as observed in work on Contrastive Sparse Representations (CSR) and CSRv2. We show that such task-specific, low-active dimension representations can be learned by training a light-weight adapter (namely, Sparsified Expert) over any base model using a novel weighted-top-k objective. Our approach departs from prior methods that either (1) concentrate information in predetermined leading dimensions (i.e., prefix-based), as in MRL and Matryoshka Adapters, or (2) project embeddings into a higher-dimensional space to identify sparse representational axes, as in CSR and CSRv2. We then introduce Mixture of Representations from Sparsified Experts (MORSE), which combines representations from multiple SEs to recover task-agnostic performance while maintaining sparsity in representations. MORSE occupies the previously underexplored regime between moderately sparse, task-agnostic representations and extremely sparse, task-specific ones, revealing a Pareto frontier between representational sparsity and generality. Across three base models, our SE outperforms other methods at ultra-low dimensions and MORSE improves over others on cross-lingual generality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.