Diffusion MoE LLMs for Memory Constrained Deployment
Abstract
Diffusion large language models (dLLMs) generate multiple tokens in parallel through iterative refinement, offering an alternative to sequential autoregressive decoding. Many recent dLLMs adopt Mixture-of-Experts (MoE) backbones, in which each token activates only a small subset of experts. However, when many tokens are processed simultaneously, their combined routes can span much the expert bank. This is particularly costly on memory-constrained devices, where the full expert bank cannot remain in DRAM and non-resident expert weights must be transferred from flash during inference. We propose dynamic just-in-time expert pruning, a training-free method that restricts routing at each layer and denoising step to a compact expert set selected from the model's current router scores. We further introduce an approach that favors experts selected in the preceding iteration, reducing expert-pool turnover and associated memory traffic. Experiments on DiffusionGemma 26B-A4B and SDAR-30B-A3B show that restricting the selected expert working set to 32 of 128 experts preserves benchmark performance within of the full baseline while substantially reducing expert-weight demand. Trace-driven simulations estimate up to relative modeled speedup on the evaluated mobile hardware configurations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.