acceptodds
Under review as a conference paper at ICLR 2027

Den2MoEE: Reconstructing Dense LLMs into Expert-Specialized Mixture-of-Experts for Efficient Embedding Models

Abstract

Most existing LLM-based embedding models in database semantic search systems rely on dense Transformer architectures. In contrast, Mixture-of-Experts (MoE) embedding models provide a cost-efficient scaling paradigm by activating only a subset of parameters. However, training large-scale MoE embedding models from scratch is prohibitively expensive. Although reconstructing dense LLM embeddings into MoE architectures offers a practical alternative, existing approaches often overlook expert specialization and routing characteristics specific to embedding tasks, resulting in homogeneous experts and inefficient performance recovery. To address these limitations, we propose Den2MoEE, a unified Dense-to-MoE reconstruction framework for embedding models that jointly promotes domain-aware expert specialization and routing-aware adaptation. Den2MoEE reconstructs dense embedding models through TokenScore-weighted neuron domain modeling, expert-diversity-driven Dense-to-MoE conversion, and an efficient two-phase recovery strategy. Furthermore, we exploit MoE router signals as complementary embedding representations, further enhancing the quality of the reconstructed embeddings without increasing activated parameters. Extensive experiments demonstrate that Den2MoEE reduces the number of activated parameters by 33%–55% compared with the original dense LLM embedding models, while consistently outperforming existing embedding baselines under comparable activated parameter budgets. Moreover, Den2MoEE delivers substantial system-level gains, achieving 1.35×–1.8× higher throughput with significantly lower inference latency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.