PRIMO: Pareto-Optimal Information-Constrained Compression for Mixture-of-Experts Language Models
Abstract
Mixture-of-Experts (MoE) architectures scale large language models (LLMs) efficiently, but every expert must be stored although only a few are active per token, so memory becomes the deployment bottleneck. Existing pipelines merge experts into a shared base by heuristics, truncate the deltas at energy thresholds that bound neither error nor cost, and deploy one fixed configuration for every task. We propose PRIMO (Pareto-optimal Information-constrained Mixture-of-experts Compression), which casts post-training MoE compression as one constrained problem: minimize model size subject to explicit bounds on reconstruction error and computational overhead. Its key insight is that the costly stages are task-independent while the right compression level is task-dependent, and routing links the two. Once per model, MI-constrained decomposition builds the shared base by minimizing a routing-weighted KL surrogate of information loss, PGD low-rank compression with a rank search fits the deltas to both bounds, and a Pareto frontier stores the non-dominated size-error configurations. At deployment, task-adaptive selection tightens the error bound by the target task's routing shift and picks a configuration by a linear scan after one forward pass, without re-optimization. On four MoE LLMs (Mixtral-87B, DeepSeek-MoE-16B, Phi-3.5-MoE, Qwen2-57B-A14B) and 12 tasks, PRIMO attains the lowest perplexity and highest average accuracy in all eight settings, and its lead over D-MoE widens with compression, reaching 5.12 against 6.46 WikiText-2 perplexity at 60% compression of Mixtral-87B. Under routing shift, task-adaptive selection raises HumanEval Pass@1 from 21.5% (static configuration) to 33.5%, and its gain grows with the shift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.