acceptodds
Under review as a conference paper at ICLR 2027

TAP-MoE: Task-Aware Profiling for Memory-Efficient Mixture-of-Experts Deployment

Abstract

Mixture-of-Experts (MoE) models scale capacity through sparse activation, yet their large expert pools impose substantial storage and weight-transfer costs that hinder efficient edge deployment. When experts cannot all reside in accelerator memory, deployment must address both compression-induced quality loss and expert-loading overhead. We propose TAP-MoE, a task-aware framework that connects fine-grained expert behavior to task-level deployment decisions. It captures expert activation and sensitivity patterns through shared class-conditional statistics. Reweighting these statistics by the target task's token-class distribution yields a workload-specific profile. This profile guides sensitivity-aware mixed-precision allocation, task-conditioned expert prefetching and residency, and risk-aware output compensation. Precision allocation determines both expert sizes and quantization residuals, linking memory planning and error correction to the same task profile. The framework separates offline calibration from online execution. It preserves the pretrained routing rule and shared experts without model retraining. Our analysis bounds profile-transfer error and derives optimality conditions for a continuous precision-allocation surrogate. We evaluate TAP-MoE on OLMoE-1B-7B, Qwen1.5-MoE-A2.7B, and DeepSeek-MoE-16B-Base across three task domains under matched storage and runtime memory budgets. Compared with BF16, TAP-MoE reduces routed-expert weight storage by 75% while maintaining near-BF16 quality. Against the strongest evaluated SOTA baseline for each metric, TAP-MoE achieves 1.0-5.0% relative reductions in perplexity and 1.25-1.75× decoding speedups.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.