QuaL-MoE: Practical Dynamic Latent Mixture-of-Experts with Quantile Balancing
Abstract
Fixed Top- routing assigns every token in a mixture-of-experts (MoE) model the same number of experts, regardless of how much computation it needs. Dynamic routers removes this restriction, but they typically control the average budget only indirectly and handle load balancing separately. We present QuaL-MoE, a dynamic MoE router that uses expert-specific quantile thresholds to determine which tokens each expert accepts. This naturally balances expert loads while keeping the average number of activated experts close to a target, yet allowing the number of experts to vary across tokens. To make this routing practical at scale, we introduce global normalization, which uses one normalized score for both expert selection and output weighting, allowing task-loss gradients to directly shape routing decisions. We also develop distributed estimators that aggregate the thresholds across training ranks. A single inference-time coefficient scales the thresholds, letting one trained model trade quality for computation without retraining. At the 1.2B scale, QuaL-MoE outperforms Top- and DynMoE while keeping the average expert count closer to its target. In 10B-parameter pretraining, it surpasses an auxiliary-loss-balanced Top- model while activating 11.8% fewer experts at inference. Our analysis shows that allocation varies across task domains, declines with token position during both prefill and decoding, and increases with the number of reasoning steps in synthetic tasks. Finally, combining QuaL-MoE with latent experts makes allocation finer-grained: QuaL-MoE-Latent outperforms an architecture-matched Top--Latent model throughout training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.