acceptodds
Under review as a conference paper at ICLR 2027

SP-MoE: Post-Training Neuron Sharing for Memory-Constrained Inference

Abstract

Mixture-of-Experts (MoE) language models activate a few experts per token, but all experts occupy storage, making deployment memory a bottleneck. Existing work addresses this at two levels. At the system level, offloading keeps a subset of experts resident and fetches the rest from host memory; under a tight budget these full-expert transfers can dominate decoding: a Mixtral-8x7B baseline spends 86.8% of decode time waiting on transfers at a 16GiB weight budget, even with prefetching. At the model level, pruning removes experts and merging consolidates them, reducing storage but potentially losing expert-specific capacity; each retained expert remains a whole transfer unit. Both levels leave the transfer unit at expert granularity. We introduce SP-MoE, a post-training neuron-sharing method that reduces the size of this unit while retaining every routed expert. It aligns intermediate neurons, selects low-sensitivity ones by reconstruction-gradient magnitude, and aggregates them into one shared component per layer, leaving a private residual per expert. The shared component stays resident, so a miss transfers only that residual. Because a SwiGLU neuron is a gate column, an up column, and a down row whose contributions add in the down projection, decomposed execution computes the shared branch once per token and combines it with routing-weighted private outputs, reproducing the transformed model exactly. Across OLMoE, Qwen1.5-MoE, and Mixtral, SP-MoE achieves the highest average zero-shot accuracy among tested compression baselines at 50% expert-parameter sparsity. On one RTX 4090 at batch size one, under per-model weight budgets including resident shared weights, it improves throughput at every tested sparsity, by up to 1.71 over HC-SMoE and 2.09 over Full MoE. On Mixtral at 50% sparsity, each miss transfers 144 rather than 336 MiB.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.