SERVAL: Efficient In-Flight Draft Adaptation for LLM Serving
Abstract
Speculative decoding accelerates large language model inference: a draft model proposes tokens and the target model verifies them. Its speedup depends strongly on the acceptance length, the number of tokens each verification step accepts, which often falls when a fixed draft encounters out-of-distribution traffic. We introduce SERVAL (SERVing-time Adaptive LoRA), a speculative-decoding serving system that adapts the draft (e.g., DFlash) to the traffic it serves using persistent per-domain LoRA adapters. At each decode step, the server sends the target model's hidden states and token-level acceptance outcomes to a trainer on separate GPUs, which updates the adapters asynchronously. The adapters are applied only to the query and output projections of a DFlash draft, whose prefix key–value (KV) cache is computed from the target's hidden states, so running requests switch to new adapter versions at the next decode step after they arrive, without recomputing that cache. Starting from a general adapter it trained earlier, SERVAL has a higher acceptance length and a lower mean time per output token than the frozen DFlash draft for three models on every out-of-distribution stream and benchmark we evaluate. On a mixed stream of code, finance, and medical requests, it raises the acceptance length by 12–57% and lowers the time per output token by 12–51%. The server's overhead for sending feedback and loading new adapter versions is only 0.6–2.1% of the decode step. The per-domain adapters it learns can be saved for reuse: served frozen by another server, they keep a higher acceptance length than DFlash and the general adapter on Qwen3-8B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.