ProFiMoE: Promotion-Aware Multi-Fidelity Inference for Mixture-of-Experts Models
Abstract
Post-routing compression of mixture-of-experts (MoE) models often saves computation by skipping selected routed experts, removing their token-specific responses. We introduce full-or-light execution: each routed expert runs either its Full Expert or a lower-cost Light Subnet. Since Light execution already preserves part of the expert response, the relevant allocation signal is the magnitude of the response recovered by upgrading an expert from Light to Full. We call this signal Promotion Utility. ProFiMoE applies this principle to frozen pretrained MoE models without retraining. It constructs contiguous Light Subnets from calibration-ranked intermediate channels and estimates Promotion Utility by combining an expert-contribution prior with a calibrated Fidelity Gap. Promotion-Aware Allocation then assigns Full execution through a cumulative Top- rule, while Observed-Mass Rescaling adjusts the mixed-fidelity output using calibrated Prefix Coverage. We evaluate three MoE models at target effective FFN pruning ratios from 10% to 60%. At 50% and 60% target pruning on Qwen3, the reported seven-task mean is 2.83 and 2.81 points above the ACE reference. For 1024-token inputs on an NVIDIA A100 at 60% target pruning, time to first token and time per output token decrease by up to 50.7% and 22.7%, respectively, relative to BF16.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.