CurveK: Throughput-Curve Estimation for Batched Speculative Decoding
Abstract
Speculative decoding accelerates large language model inference by drafting multiple tokens and verifying them simultaneously, but its efficiency critically depends on selecting an appropriate draft length \(k\). This choice is particularly challenging in batched serving environments, where diverse requests share a common draft length but exhibit different acceptance patterns. Existing methods dynamically adjust \(k\) using arm-level rewards, draft-side statistics, or acceptance estimates. However, we theoretically show that these signals remain coarse: they do not directly represent how acceptance changes across draft positions in the current batch. This position-specific view defines a throughput curve over candidate draft lengths, whose maximum gives the draft length selected by our objective. Motivated by this insight, we formulate dynamic draft-length selection as estimating the batch throughput curve and choosing its peak. We propose CurveK, a lightweight and easy-to-implement method that estimates this curve from accepted-token depths already returned by the verifier. Importantly, CurveK requires neither additional training nor changes to the verification process. Experiments with 8B and 70B models in vLLM show throughput improvements over baselines across diverse workloads, including mixed-workload settings. Ablation studies validate the design of CurveK, and further analysis shows that its curve estimates consistently select high-throughput draft lengths.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.