TraceDraft: Online Draft Adaptation During Speculative Decoding
Abstract
Speculative decoding, which verifies draft tokens in a single target-model pass, has been widely adopted as an effective method for accelerating language model inference. Online distillation has been used to adapt speculative drafts to the requests encountered during deployment. However, methods that train on completed requests or request batches defer the benefit of their feedback to subsequent requests. Consequently, exploiting token-level feedback before the current response ends, which enables immediate adaptation from ongoing inference without first generating complete responses for training, remains a challenge. As real request streams often span multiple domains, retaining useful adaptation across domain changes is also worth addressing. To tackle these challenges, we propose TraceDraft, a lightweight framework for online draft adaptation within responses. TraceDraft uses target predictions on the prompt and verified continuation as supervision, updating only the draft's output normalization and prediction head. Two training modes are available to deal with different situations: Reset mode confines adaptation to each conversation, while Persistent mode retains updates across requests with a built-in anchor toward the initial weights. Updates of TraceDraft remain separate from outer training, enabling combination with other online learners. We evaluate eight workloads, including AIME2024/2025 and MT-Bench. Experimental results show gains in accepted length and end-to-end throughput on all eight workloads, with long-output gains on DeepSeek-R1-Distill-Llama-8B averaging 10.84% and 9.16%, respectively, over the fixed draft. Combining TraceDraft with online learners also adds up to 10.03% throughput over their respective baselines. In cross-domain experiments, Persistent improves throughput by up to 21.11% over the full stream and up to 12.10% after the domain switch relative to the fixed draft. These results suggest that TraceDraft provides a practical framework for learning from incoming requests to improve speculative decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.