REPS: Compressed Expert Replacement for Multiversion MoE Serving
Abstract
Serving expert-local customizations of a mixture-of-experts (MoE) model requires a representation that supports both local updates and shared execution. We propose Replaceable Expert Patches for Serving (REPS), which uses the same compressed records for a parent and its complete expert replacements. Two sign bytes encode each width-eight block through a complete Hadamard transform. This structure gives independent scalar assignment at fixed scale and integer reconstruction without codeword lookup. A version-aware operator selects parent or patch rows and consumes these records directly during prefill. Qwen3-30B-A3B Code updates improve target perplexity while keeping non-target and general-text degradation within 1% in three fixed-initialization data orders. On one of these exact updates, integer reconstruction reduces synchronized backend first-token time by 3.51 and 4.45 at batches one and eight relative to our lookup-table (LUT) decoder, with unchanged weights and decode. In a nine-format storage control, sharing reduces REPS weight allocation by 48.9% and 73.4% at two and four structural versions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.