Activation Factorization for Dense-to-MoE Conversion
Abstract
Converting pretrained dense language models into Mixture-of-Experts (MoE) architectures can enable structured sparse execution without pretraining a new model from scratch. The quality of the converted model, however, depends on how neuron activation structure is represented for expert construction and rout- ing. We propose a dense-to-MoE conversion framework based on non-negative matrix factorization (NMF) of FFN activation magnitudes. The factorization provides complementary neuron-side factor loadings and token-side factor co- efficients. We use the neuron-side representation to construct balanced physi- cal experts, while a lightweight predictor estimates token coefficients and maps them to expert-level routing scores through the same factorized activation struc- ture. The original dense-model parameters remain frozen, with only lightweight routing modules adapted after conversion. On LLaMA-3-8B with 25% active FFN capacity, our method achieves an average accuracy of 57.93% across five downstream benchmarks, exceeding the average accuracy of all evaluated CMoE configurations. Controlled ablations show that both NMF-based expert construc- tion and NMF-guided routing contribute to the final performance. On a single NVIDIA RTX 4090, the resulting physical expert execution achieves 1.13–1.22× end-to-end prefill speedup across sequence lengths from 512 to 2048. These re- sults show that activation factorization provides an effective representation for low-cost dense-to-MoE conversion with structured sparse execution
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.