acceptodds
Under review as a conference paper at ICLR 2027

BMoE: By segmenting experts to share and amplify single-layer routing candidates and reducing the total parameters

Abstract

Sparse mixture-of-experts (MoE) models activate few experts per token, yet must store the entire expert bank. Cross-depth sharing reduces this storage, but also asks shared experts to serve representations from different layers. We propose BMoE, a block-wise MoE that separates shared expert capacity from layer-specific transformations. Contiguous layers reuse a pool of latent expert cores, while learned input/output interfaces, routers, and attention remain layer-specific. Each depth can therefore access the same nonlinear cores through different learned maps. We evaluate eight 36-effective-layer configurations trained for approximately 208B scheduled token positions with matched selected-expert computation. At nearly equal total parameters (7.063B versus 7.056B), two-layer BMoE reduces WikiText-103 perplexity from 17.867 to 17.031 and raises the nine-task mean from 50.74 to 52.10 relative to a per-layer latent MoE. Four-layer BMoE improves all four aggregate metrics with 43.2% fewer parameters. It also outperforms whole-layer looping at matched expert storage. Tying the latent interfaces saves only 1.6% of parameters but degrades every aggregate metric; forward diagnostics show differentiated routing within shared pools. Together, these results show that the boundary of depth sharing matters: expert cores can be efficiently reused across depth while layer-specific transformations preserve flexible access to the shared capacity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.