acceptodds
Under review as a conference paper at ICLR 2027

Expert-Level Analog-Digital Mapping of Mixture-of-Experts with Theoretical Generalization Guarantees

Abstract

Mixture-of-Experts (MoE) models enable scalable learning by activating only a subset of parameters per input, but remain memory-bound due to frequent movement of expert weights. Analog in-memory computing (AIMC) offers a promising solution by performing matrix–vector multiplications (MVMs) directly within memory, reducing data movement and improving energy efficiency. However, existing AIMC deployments map all MVMs uniformly to the analog units, leading to significant accuracy degradation for large MoE models due to hardware non-idealities. In this paper, we show that noise sensitivity varies significantly across MVM operations, with dense modules (such as attention layers) and certain experts being particularly vulnerable to analog noise. We then introduce a theoretically grounded metric, the maximum neuron norm score (MaxNNScore), to identify noise-sensitive experts. Our theoretical results prove that selectively mapping high-MaxNNScore experts to the digital units significantly improves the tolerable noise level while maintaining generalization compared to uniform analog execution. Empirical evaluations on large-scale MoE models, including DeepSeekMoE (16B), Qwen-1.5 MoE (14B), and OLMoE (7B), demonstrate that our approach recovers near full-precision performance (within 3-8% drop) over several benchmark LLM tasks while mapping only 11-29% of parameters to digital.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.