CH-MoE: Decoupling Learning Credit from Hard Execution Budgets
Abstract
Sparse mixture-of-experts (MoE) routing often struggles to balance dynamic expert selection for model quality with strict hardware execution budgets for system efficiency. We introduce CH-MoE, a unified routing framework that decouples hard execution budgets from gradient-based learning credit. At the Llama 260M/5B scale, under a strict budget of 80B expert calls, token-local CH-Global improves validation perplexity and macro accuracy over ReMoE and Default MoE while driving substantial load variance reductions (dropping Load CV by over 85%). Rigorous ablation controls demonstrate a structural proof-of-concept: our mask-closed architecture maintains robust language generalization without relying on conventional confidence penalties or proxy regularizations. At the 4.38B scale, our full-sequence allocator (CH-Primary) translates algorithmic frugality into measured hardware speedups, increasing single-A100 throughput by up to 26.03%. Crucially, CH-Primary outperforms a structurally matched same-work scalar baseline by 11.22%, indicating that its distinct allocation distribution provides improved execution efficiency even under identical load-balancing constraints. Furthermore, physical two-node (8A100) execution delivers up to 49.95% higher throughput, driven by a 63.41%–68.32% collapse in receive-route variance. Together, these results establish CH-MoE as a rigorously controlled framework for explicitly uncoupling execution limits from sparse learning signals.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.