SpecSys: Selective Speculative Decoding and Completion-Aware Scheduling for Efficient Agentic LLM Inference
Abstract
Large language model (LLM)-powered agents often invoke multiple LLM requests, making LLM inference a major contributor to end-to-end task completion time. Our measurements reveal that the dominant bottleneck shifts with serving load: from decoding at low concurrency to queueing at high concurrency. Speculative decoding (SD) is a state-of-the-art technique for accelerating LLM inference, but its benefits diminish as serving load increases and can even become negative under high load. Existing adaptive SD methods cannot eliminate this problem because they lack native non-SD execution, while conventional serving schedulers do not fully exploit workflow structure and completion progress of agent tasks. To address these challenges, we design SpecSys, an efficient inference system for agent workloads that enables load-aware selective SD with completion-aware scheduling. SpecSys exposes an externally controllable runtime interface that enables complete iteration-level switching between native SD and non-SD execution. It also employs a lightweight history-based estimator to predict the remaining steps of each agent task based on its workflow state and request characteristics. SpecSys adapts these mechanisms with serving load: it selectively enables SD to reduce decoding latency under low load, while prioritizing tasks with fewer predicted remaining steps under high load to reduce queueing delay. Across three LLMs, four agent workloads, three agent workflows, and two hardware platforms, SpecSys consistently reduces average job completion time (JCT) over strong baselines, including ThunderAgent, vLLM Dynamic SD, and SGLang Adaptive SD, achieving up to a 2.10 speedup under resource-constrained settings. The source code of SpecSys is anonymously available at https://anonymous.4open.science/r/SpecSys-Code-E03D/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.