Queue-Aware Combinatorial Bandit Scheduling for Speculative Decoding in Cloud–Edge LLM Serving
Abstract
Speculative decoding accelerates autoregressive LLM inference by enabling a draft model to generate candidate token blocks, which the target model then verifies in parallel. Existing research largely targets draft-model quality and point-to-point cloud-edge efficiency, leaving unaddressed the contention among edge-generated verification demands for the central target model's resources under multi-user workloads. The scheduling problem that arises when candidate blocks generated asynchronously by multiple edge nodes compete for limited verification resources has become a major bottleneck in improving system efficiency. To address this problem, we formulate shared-verifier admission as a queue-aware combinatorial semi-bandit problem for speculative decoding and propose the Speculative Decoding Pipeline Scheduling (SDPS) algorithm. Each node places its locally generated draft blocks into a pending-verification queue, while the central verifier selects a set of nodes in each round based on queue pressure, verification yield, and latency constraints. Without prior knowledge of node-specific yields, SDPS provably reduces its utility loss as the scheduling horizon grows, while average pending backlog is bounded under a windowed load condition through a single control parameter. Experiments on a distributed LLM system demonstrate that SDPS effectively improves the utilization of centralized verification resources while controlling pending backlog and protecting latency SLOs under multi-user contention. Code is available at https://anonymous.4open.science/r/SDPS1-2E52/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.