acceptodds
Under review as a conference paper at ICLR 2027

Tying the Loop: Tied Expert Layers in Mixture-of-Experts Language Models

Abstract

Mixture-of-Experts (MoE) architectures efficiently scale Large Language Models (LLMs) by activating only a small fraction of their experts per token, yet the full parameter count—dominated by the expert parameters—must be held in training and inference memory. To address this, we systematically study expert tying, an architectural modification that shares expert parameters across consecutive transformer layers while preserving independent, layer-wise routing and attention. We evaluate this approach across common, state-of-the-art architectures, including OLMoE, Qwen3, and DeepSeek-style MoEs. On full-size OLMoE-1B-7B, tying experts halves unique weight storage with virtually unchanged validation loss and similar downstream accuracy after 15.7B training tokens. Controlled ablations identify the value of preserving layer-specific attention under expert sharing. Reinvesting saved parameters in larger shared expert pools lowers validation loss at the baseline parameter budget across all configurations, making expert tying useful for both reducing storage and improving parameter allocation, and providing a more favorable compute-to-memory trade-off.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.