LSV: EFFICIENT REVERSIBLE GPU MULTIPLEXING FOR LLM SERVING
Abstract
Intra-GPU multiplexing lets Preffll and Decode share one LLM replica while using different GPU resources. The useful split changes with the request mix. Existing controllers can select a new split quickly. But work already submitted under the old split can delay the ffrst useful work under the new split by milliseconds. We call this delay reconffguration debt. We present Logical Stream Virtualization (LSV), an execution layer with stable logical lanes for Preffll and Decode. LSV tracks issued and completed work with resource epochs. It limits future old-epoch commitments and changes both lanes at one safe boundary. We implement LSV in SGLang and evaluate it on an eight-H100 NVLink server across three model scales and two model families. A controlled test shows that host handoff stays near 9.5 µs, while physical activation reaches 16.3 ms after sixteen old-epoch replays. Across the three models, LSV preserves more latency headroom, compliant goodput, and joint SLO attainment as load rises. The results show that adaptive multiplexing needs a runtime contract for when a resource decision becomes physical, not only a policy for which split to choose.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.