Agent Turns Are Not Chat Turns: Matching Draft Supply with Verification Cost
Abstract
Speculative Decoding (SD) accelerates LLM serving by verifying several draft tokens in one forward pass. Current SD approaches work well for conversation workloads, where contexts are short and content is largely new. Yet agent workloads are different: a coding agent solves tasks through tens of turns of generation, tool use and observation, building contexts that reach tens of thousands of tokens. Much of what the agent writes—re-edited files, re-issued commands, summarized context—closely follows text already in the prompt: 70% of generated tokens match the agent's own earlier output, and the reuse grows with the session. The prompt itself thus provides a growing candidate supply that needs no auxiliary model, but does need a retrieval-based drafter that can adapt it. Scaling retrieval-based drafting to concurrent serving raises two problems: building a per-request index blocks the scheduler thread, and the draft tree must be sized to the verification capacity the hardware affords. We present THRIFT, a training-free retrieval-based drafter that indexes the prompt with a suffix automaton and the current generation with a trie, builds each request's corpus asynchronously, and sizes the draft tree to the batch's verification capacity. We also propose an evaluation protocol that stratifies turns by their relation to the agent's earlier output and reports decode-stage next to end-to-end speedup. On BeyondSWE with Qwen3-14B at batch size 8, THRIFT achieves the best end-to-end speedup up to over autoregressive decoding; its accept length grows from to tokens per forward across a session while training-based baselines stay flat.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.