acceptodds
Under review as a conference paper at ICLR 2027

DELMOE: DUAL-SPARSE ROUTE DELETION WITH RECTIFICATION FOR EFFICIENT BATCHED DECODING

Abstract

Mixture-of-experts (MoE) models expand model capacity through sparse expert activation without proportionally increasing per-token computation. In batched decoding, different requests can collectively activate many experts, increasing latency without necessarily improving accuracy. Yet skipping experts can reduce accuracy, and surviving gate weights may no longer suit the reduced expert mixture. Existing batch-aware route replacement reduces active experts but largely preserves route counts and redistributes weights without calibrating them for accuracy under the altered routing state. We propose DelMoE with two modules for retraining-free batched MoE decoding. (1) Reuse-aware dual-sparse route deletion: We build a backbone from high-priority native routes, reuse its experts across requests, and supplement the retained set. Deleting routes outside this set reduces both active experts and token–expert routes. (2) Routing-aware gate-weight rectification: We calibrate one rectification coefficient per model offline and keep it fixed during inference to adjust surviving gate weights. To avoid executing deleted zero-weight slots, our custom vLLM CUDA backend processes retained routes as variable-length workloads. Across three MoE models and eight benchmarks, DelMoE achieves higher average accuracy and generally faster decoding than batch-aware route-replacement baselines, with particularly large accuracy gains at high sparsity. Compared with Vanilla, DelMoE achieves up to 2.49× decoding speedup at comparable or higher average accuracy while activating fewer experts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.