acceptodds
Under review as a conference paper at ICLR 2027

COMPRESS FOR RECOVERABILITY: STREAMED COUPLED CORRECTIONS FOR MOE EXPERTS

Abstract

Mixture-of-Experts (MoE) models activate few experts per token, but their expert banks consume substantial GPU memory. We study recoverability as a third axis for compression, alongside resident memory and uncorrected accuracy: how much lost computation can be restored under a limited correction budget? A compressed model with poor standalone accuracy can still provide an effective base if its errors admit sparse, input-dependent correction. This perspective makes expert merging a viable compression strategy when combined with selective recovery. We introduce Delta-MoE, which keeps compressed experts on the GPU and streams selected residual channels from host memory. Each correction couples the gate, up, and down projections of a SwiGLU channel and recomputes its nonlinear contribution. Our analysis characterizes recovery through residual error spectra and explains why improving an uncorrected merge need not improve its corrected counterpart. On Qwen3-30B-A3B, the fused implementation achieves 0.810 GSM8K accuracy at 16× resident expert compression, with MATH-500 and IFEval point estimates similar to an 8× quantized base, and Mixtral-8x7B also exhibits recovery from merging.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.