ShadowMoE: Graph-Conditioned Shadow Expert Representations for Mixed-Precision MoE Quantization
Abstract
Sparse mixture-of-experts (MoE) models expose a structured trace of computation through token-level routing: experts are not used independently, but repeatedly cooperate across layers. We ask whether this routing topology contains information about expert vulnerability to model compression. We introduce ShadowMoE, a topology-aware representation framework that converts token-level top-k routing traces into weighted cross-layer expert graphs and learns relational representations of individual experts. A graph-restricted encoder propagates information only along observed routing paths, while leakage-controlled structural objectives test whether the learned representation captures non-local routing information. We further probe whether these representations predict expert-level quantization residuals and use them to construct budget-constrained mixed-precision policies. On Qwen3-30B-A3B and Mixtual V0.1, ShadowMoE achieves a strict link-prediction AUC of 0.9432, compared with 0.8378 for a graph-free MLP, and its expert representations explain R2 = 0.5405 and 0.5953 of FP8 and INT4 local quantization residuals, respectively. At an average 5.0-bit budget, the resulting INT4/FP8 allocation retains 99.43% and 98.33% of FP8+activation quantized accuracy on the two models, respectively. Importantly, independently estimated expert risks do not reliably compose into scheme-level degradation: the correlation between predicted local risk and complete-scheme loss is only. 0.0580. This negative result suggests that compression sensitivity in MoEs is relational rather than purely expert-wise. ShadowMoE therefore provides both a practical compression mechanism and a framework for studying how routing topology encodes functional dependencies among experts. Anonymous artifacts: https://anonymous.4open.science/r/topology_0723-70B3.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.