acceptodds
Under review as a conference paper at ICLR 2027

SERVAL: Efficient In-Flight Draft Adaptation for LLM Serving

Abstract

Speculative decoding accelerates large language model inference: a draft model proposes tokens and the target model verifies them. Its speedup depends strongly on the acceptance length, the number of tokens each verification step accepts, which often falls when a fixed draft encounters out-of-distribution traffic. We introduce SERVAL (SERVing-time Adaptive LoRA), a speculative-decoding serving system that adapts the draft (e.g., DFlash) to the traffic it serves using persistent per-domain LoRA adapters. At each decode step, the server sends the target model's hidden states and token-level acceptance outcomes to a trainer on separate GPUs, which updates the adapters asynchronously. The adapters are applied only to the query and output projections of a DFlash draft, whose prefix key–value (KV) cache is computed from the target's hidden states, so running requests switch to new adapter versions at the next decode step after they arrive, without recomputing that cache. Starting from a general adapter it trained earlier, SERVAL has a higher acceptance length and a lower mean time per output token than the frozen DFlash draft for three models on every out-of-distribution stream and benchmark we evaluate. On a mixed stream of code, finance, and medical requests, it raises the acceptance length by 12–57% and lowers the time per output token by 12–51%. The server's overhead for sending feedback and loading new adapter versions is only 0.6–2.1% of the decode step. The per-domain adapters it learns can be saved for reuse: served frozen by another server, they keep a higher acceptance length than DFlash and the general adapter on Qwen3-8B.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.