acceptodds
Under review as a conference paper at ICLR 2027

LSV: EFFICIENT REVERSIBLE GPU MULTIPLEXING FOR LLM SERVING

Abstract

Intra-GPU multiplexing lets Preffll and Decode share one LLM replica while using different GPU resources. The useful split changes with the request mix. Existing controllers can select a new split quickly. But work already submitted under the old split can delay the ffrst useful work under the new split by milliseconds. We call this delay reconffguration debt. We present Logical Stream Virtualization (LSV), an execution layer with stable logical lanes for Preffll and Decode. LSV tracks issued and completed work with resource epochs. It limits future old-epoch commitments and changes both lanes at one safe boundary. We implement LSV in SGLang and evaluate it on an eight-H100 NVLink server across three model scales and two model families. A controlled test shows that host handoff stays near 9.5 µs, while physical activation reaches 16.3 ms after sixteen old-epoch replays. Across the three models, LSV preserves more latency headroom, compliant goodput, and joint SLO attainment as load rises. The results show that adaptive multiplexing needs a runtime contract for when a resource decision becomes physical, not only a policy for which split to choose.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.