acceptodds
Under review as a conference paper at ICLR 2027

Cadence: Regulating the Speculation Externality in LLM Serving

Abstract

Speculative decoding (SD) speeds up LLM inference by verifying draft tokens in parallel. SD applies to every request, but its effectiveness, how many draft tokens are accepted, varies across requests and over time. Requests in a batch share each step, so chasing mean throughput can slow requests that accept little: in one step at batch 112, SGLang's default SD (SGL-SD) raises throughput 13% while 26% of requests run slower than autoregressive (AR) decoding. Cadence, our per-step controller, caps each request's delay, its time lost relative to AR steps. Every step commits at least one token per request, so the caps reduce to one allowed step latency, known before drafting. Within this latency, an all-or-none floor gives every request a draft token whenever the latency allows, and Cadence picks the verification budget that an operator's utility scores highest. On Qwen3-32B under high load, Cadence improves goodput by 1.29 over AR decoding and by 1.80 over SGL-SD, while 99.5% of its requests run no slower than AR decoding, against 69% under SGL-SD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.