Unmasking the Hidden Bias, Fairness, and Safety Costs of Compression with Mixture-of-Expert Models
Abstract
Mixture-of-Experts (MoE) language models are often compressed for deployment, yet the fairness, bias, and safety effects of these interventions remain poorly understood, particularly for Expert Compression (EC) algorithms and their required calibration data. We investigate how compression algorithms and calibration data mixtures affect fairness and safety behaviour in MoE language models. Across three diverse MoE architectures, we find that EC is highly sensitive to calibration data composition. Expert merging produces the largest regressions on average, although the relative behaviour of merging algorithms varies substantially across architectures. Expert pruning is generally less disruptive, while quantization and weight sparsity are more stable and less sensitive to calibration data. We show that benchmark level improvements can be misleading: some aggressive compression settings appear more stable on selected fairness benchmarks while degrading task accuracy and instruction following, and prompt matched generation diagnostics reveal greater behavioural drift from the uncompressed baseline. Targeted interventions show that pruning just three of the 6,144 experts in Qwen3-30B-A3B can turn a refusal of a harmful AdvBench request into compliance, whereas quantizing the same experts to 2 bits preserves the refusal. These results demonstrate that the fairness and safety effects of MoE compression are neither simple nor monotonic and highlight calibration data, model architecture, and algorithm selection as key considerations when compressing MoEs for deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.