acceptodds
Under review as a conference paper at ICLR 2027

ReFrame: Elastic Expert–KV Memory Sharing for Mixture-of-Experts Serving

Abstract

In large mixture-of-experts serving systems with limited GPU memory, expert weights and key–value (KV) state compete for capacity, yet most existing systems statically partition this memory between them. Since only a subset of resident experts is activated per token, substantial expert memory can remain underutilized. Therefore, static GPU memory partitioning prevents the underutilized memory from being used by the KV cache, causing requests to wait for KV capacity and reducing overall memory efficiency. To address this limitation, we present ReFrame, a serving runtime that dynamically shares GPU memory between expert weights and KV cache while preserving sufficient expert residency. First, typed frame rebinding enables GPU memory to alternate between expert and KV cache while retaining the native tensor and block views required by each consumer. Second, a bounded allocation policy coordinates when and how much memory to borrow, confines each loan to the triggering request, and uses private capacity first to leave shared blocks available for other requests. Finally, we evaluate ReFrame on two mixed workloads with Qwen3-235B-A22B. Against CPU-offloading baselines, the complete system achieves geometric mean speedups of 2.75× over vLLM 0.29 and 2.94× over SGLang 0.5.17 in trace completion time. An independent matched-batch comparison with vLLM 0.29 yields 2.74×, and a separate comparison with its 2,048-token batch budget yields 3.15×.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.