Cadence: Regulating the Speculation Externality in LLM Serving
Abstract
Speculative decoding (SD) speeds up LLM inference by verifying draft tokens in parallel. SD applies to every request, but its effectiveness, how many draft tokens are accepted, varies across requests and over time. Requests in a batch share each step, so chasing mean throughput can slow requests that accept little: in one step at batch 112, SGLang's default SD (SGL-SD) raises throughput 13% while 26% of requests run slower than autoregressive (AR) decoding. Cadence, our per-step controller, caps each request's delay, its time lost relative to AR steps. Every step commits at least one token per request, so the caps reduce to one allowed step latency, known before drafting. Within this latency, an all-or-none floor gives every request a draft token whenever the latency allows, and Cadence picks the verification budget that an operator's utility scores highest. On Qwen3-32B under high load, Cadence improves goodput by 1.29 over AR decoding and by 1.80 over SGL-SD, while 99.5% of its requests run no slower than AR decoding, against 69% under SGL-SD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.