acceptodds
Under review as a conference paper at ICLR 2027

Tiered Expert Embedding: Input-Side Expert Embeddings for Tiered Compression of Mixture-of-Experts LLMs

Abstract

Mixture-of-Experts (MoE) architectures scale model capacity without scaling per-token compute, but their massive expert parameters dominate memory and storage cost. Merging each layer’s experts is the direct way to shrink them: several original experts are merged into one and share its intermediate dimension. This remains viable because of a general property of gated FFNs: for any given input, only about half of an expert’s intermediate units pass the gating. The active subnetworks of merged members therefore rarely saturate the shared weights, and the residual merge error stays low-dimensional enough to be cancelled by a per-slot shift of the expert’s input. We propose Tiered Expert Embedding (TEE), a simple compression framework that (i) merges each layer’s experts with tiered, frequency-aware compression ratios, (ii) freezes the router and remaps original expert slots onto merged experts, and (iii) attaches a tiny per-slot expert embedding w0 on the input side to compensate for merging error. On Qwen3.6-35B-A3B, TEE compresses expert parameters by 5.33× and total parameters by 4.07× while leaving per-token activated parameters unchanged. Even with random expert pairing at a uniform 2:1 ratio, the pipeline loses only 0.8 points on MMLU-Pro: at 2× the capacity slack needs no pairing structure at all. At the full 5.33× ratio the compressed model is comparable to the purpose-built Qwen3.5-9B dense model on MMLU-Pro (69.3 vs. 69.2) while activating 2.6× fewer parameters per token, and outperforms MoLA under matched compression by +8.7 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.