SpecRoute: Multi-Step Speculation for Efficient LLM Agent Serving
Abstract
LLM agents interleave model reasoning with tool execution, forming observation-dependent sequences of model calls that incur substantial end-to-end latency. Conventional serving optimizations focus on individual LLM calls, while single-step action speculation overlaps the next tool invocation with ongoing reasoning. Both leave the agent’s multi-turn execution chain largely on the critical path. We present SpecRoute, a serving runtime that accelerates LLM agents through multi-step speculative execution. A lightweight draft agent speculates on future tool actions, interleaving tool execution with draft-model inference while the target processes its current context. At target tool-call boundaries, the target performs incremental, prefill-based verification of speculative observations and incorporates accepted results with KV reuse. A slack-aware scheduler coordinates batch formation and target–draft engine switching on a shared GPU, balancing speculative progress against interference with target execution. Across six agent benchmarks, SpecRoute achieves 1.28–1.56× single-request speedups over vLLM with comparable task performance. Under online serving, mean completion time decreases by 28–36% at low load and 9–20% at the highest tested loads.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.