acceptodds
Under review as a conference paper at ICLR 2027

HARP: Hierarchical Adaptive Redistribution for Pruning Mixture-of-Experts Models

Abstract

Channel-level pruning can reduce both the parameter footprint and activated computation of mixture-of-experts (MoE) models, but doing so effectively requires deciding how a fixed global channel budget should be distributed across layers and experts, as well as which channels should be retained. Existing approaches often rely on calibration activations, whose quality can vary substantially across routed experts and calibration distributions. We present HARP, a calibration-free framework for hierarchical adaptive redistribution of MoE channel capacity. HARP derives a model-intrinsic structural preservation (SP) prior from the relative fluctuation of retained absolute-weight mass under random thinning and applies the same functional at the layer, expert, and coupled-channel levels. A hierarchical allocator uses these priorities to construct hardware-aligned heterogeneous expert widths under an exact global budget, while Channel-SP realizes each width through coupled channel selection. When calibration data are available, empirical criteria can refine channel identities without changing the global allocation. Across Qwen3, Qwen3.6, and DeepSeek-V2-Lite at 25% and 50% channel pruning, HARP achieves the highest mean retained performance among the evaluated channel-pruning methods in all six settings, consistently outperforming calibration-based Wanda and ENP. SP anchoring improves calibration-assisted channel selection and reduces calibration sensitivity. Finally, a width-aware vLLM backend executes heterogeneous experts at their retained widths, achieving – end-to-end inference speedup on a single A100 GPU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.