BitLift: Unified Mixed-Precision Quantization and Caching for Efficient MoE Inference
Abstract
Mixture-of-Experts (MoE) models reduce computation through sparse activation, but moving expert weights can dominate decoding when the model exceeds device memory. Dynamic mixed precision and DRAM-caching can reduce this cost, yet their effectiveness depends on how weight representation, precision allocation, and cache residency interact. We introduce , a training-free framework that co-designs these components. A router-rank policy dynamically allocates a fixed nominal precision budget across selected experts. Each expert is encoded as a 2-bit base and two 1-bit refinements, jointly optimized for the target policy through weighted scale fitting and error compensation. A two-pool cache separates base and refinement planes and applies cache-resident precision refinement without fetching additional weights for an upgrade. Across three MoE model families, rank-assigned execution approaches independently optimized 3-bit quality at nominal budgets of 2.333–2.375 bits per routing slot, and outperforms the evaluated 2.5-bit baselines. On Qwen3-30B-A3B, cache-resident refinement matches uniform 4-bit WikiText-2 perplexity with a cache holding 50% of the 4-bit routed-expert footprint. At smaller matched capacities of 6.25–25%, it improves on uniform 3-bit perplexity. In this flash-limited regime, our analytical hardware model estimates a throughput gain of – over uniform 3-bit and – over uniform 4-bit inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.