acceptodds
Under review as a conference paper at ICLR 2027

RAZOR: Pruning Replaceable Experts in LLMs

Abstract

Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Pruning aims to reduce this burden at a fixed budget while preserving output distributions and, for reasoning models, reasoning ability. Deletion damage depends on functional replaceability by surviving computation, not usage or contribution magnitude alone. We introduce RAZOR, a training-free method that scores functional replaceability using consensus residuals, the deviations of expert outputs from the original weighted mixture. At a fixed layer input, an exact single-deletion identity accounts for survivor renormalization and router-selected refill. We aggregate these local scores over calibration tokens for budgeted pruning without gradients or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25% and 50% expert removal, RAZOR achieves the highest nine-task macro average among the evaluated pruning methods in all eight settings. On the two backbones with matched REAP benchmark runs, it exceeds REAP by 2.12–5.59 points and wins all 36 paired task comparisons. It also lowers reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model–budget settings. Responses generated by Qwen3.6-35B-A3B nevertheless change in diversity, formatting, and termination, showing that task retention and predictive fidelity do not ensure generation stability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.