CLASP: Cross-Layer Active-Subspace Pooling for Quantized Mixture-of-Experts
Abstract
Low-rank corrections can recover accuracy lost to aggressive post-training quantization of mixture-of-experts (MoE) language models. Most low-bit MoE correction methods fit their factors independently or within individual layers. We find substantial overlap between experts' activation-weighted dominant subspaces across layers. We use this overlap in Cross-Layer Active-Subspace Pooling (CLASP), which clusters experts by the Grassmannian distance between their active subspaces and fits a shared low-rank basis to each cluster. Sharing reduces basis storage and allows a larger correction rank at a fixed bit budget. In the evaluated Qwen models, experts activated by the same token also tend to belong to a few clusters. CLASP stores the shared factors in a device-resident pool and batches correction operations, with stream overlap reducing execution overhead. At 2.15–2.21 bits per quantized parameter, CLASP achieves the lowest perplexity among the compared quantizers on four MoE models. Under the same measurement protocol and hardware, end-to-end generation is up to faster than the FP16 baseline. Our source code will be available after acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.