Reconstructing the Routing Space: Retraining-Free Compression of Vision-Language Mixture-of-Experts
Abstract
Vision-language Mixture-of-Experts models (VLM-MoEs) hold most of their parameters in expert weights, yet their expert pools are hard to compress because visual tokens use experts far more diffusely than text tokens. Our experiments show that covering of visual routing mass takes up to as many experts as for text, and that at expert retention frequency-based pruning keeps only – of dense performance. This exposes a limitation of existing compressors: they rank experts alone or fuse parameters, ignoring the function expressiveness of the routing space. We propose \method, which selects a well-conditioned expert subset by pivoting on the eigenbasis of a routing-aware expert-output Gram matrix and, from the same basis, reconstructs the discarded routing coordinates in closed form. The procedure is one-shot and retraining-free, with no learned parameters. On Kimi-VL-A3B and Qwen3-VL-30B-A3B, \method delivers the strongest overall performance across compression budgets, and its advantage widens as compression becomes more aggressive, reaching up to points over the strongest baseline at expert retention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.