BOLD-MoE: Base-Optimized Learning of Orthogonal Dictionaries for Mixture-of-Experts Compression
Abstract
Recent Mixture-of-Experts (MoE) compression methods exploit inter-expert redundancy by representing each expert as a shared base plus a compressed expert-specific residual. However, the base is typically constructed before residual compression, ignoring that its optimal choice depends on how well the resulting residuals can be represented under a limited budget. We formulate shared-base MoE compression as a coupled optimization problem and solve it by alternating between residual compression and an activation-aware base update. This formulation is compressor-agnostic: even when retaining SVD residuals, alternating base–residual optimization improves over the one-shot decomposition. We further introduce orthogonal sparse dictionary residuals, which replace the single low-rank subspace imposed by SVD with a richer union-of-subspaces representation. Across three MoE architectures and 20-60% compression, BOLD-MoE consistently improves perplexity and zero-shot accuracy over the strongest shared-base baseline, with substantial gains persisting under strong compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.