Phase-MoE: A Co-Design to Bound the Expert Explosion in Block Diffusion Language Models
Abstract
Block Diffusion Language Models (BDLMs) have emerged as a compelling paradigm bridging autoregressive modeling and discrete diffusion, pairing inter-block causal generation with intra-block bidirectional parallel decoding. To scale these architectures efficiently, integrating Mixture-of-Experts (MoE) is essential. However, combining MoE with block-parallel decoding triggers a systemic Expert Explosion: processing multiple tokens simultaneously causes their independently routed experts to span the expert pool. This inflates memory traffic from High-Bandwidth Memory (HBM) to SRAM, creating a memory-bound bottleneck. We demonstrate that existing post-hoc mitigations break down under production continuous batching, where the temporal and semantic heterogeneity of asynchronous requests escalates global memory traffic to near-dense levels. To resolve this, we propose a bandwidth-aware co-design comprising Phase-Constrained MoE and a Phase-Aware Scheduler. Recognizing that discrete diffusion progresses through predictable mask-density phases, we condition the router's expert pools on the temporal denoising phase during training, leaving compute-bound prefill unconstrained. The Phase-Aware Scheduler subsequently clusters asynchronous requests by mask density at runtime, promoting overlapping expert execution. Under production-scale continuous batching simulations where existing mitigations fail, our co-design reduces unique expert activations by up to 56.2% and cuts MoE kernel latency by up to 38.5%, preserving competitive generative performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.