Beyond the Target: From Imitation to Collaboration in Speculative Decoding
Abstract
Speculative decoding accelerates large language model inference by letting a small draft model propose tokens for parallel verification by a larger target. For tasks evaluated by answer correctness, however, agreement with the target is an imperfect criterion: a weaker draft can occasionally offer a continuation that leads to a better final answer. Exploiting these disagreements requires learning from delayed outcomes while retaining the efficiency of parallel verification. We present Collaborative Speculative Decoding (CoSpec), a framework that learns arbitration within speculative decoding from the outcomes of complete rollouts. An arbitrator jointly observes the context, draft block, and target verification tokens through a hybrid attention mask, and decides whether to retain the draft or fall back to the target at each mismatch. Reinforcement learning combines final-answer correctness with accepted-length shaping to train these decisions. CoSpec retains parallel target verification while optimizing task utility rather than preserving the target distribution. We evaluate two model families on reasoning, code, and knowledge benchmarks. With LLaMA-3.3-70B as the target, CoSpec achieves a mean speedup of 3.96× over target-only decoding and raises the mean score on GSM8K, HumanEval, and MBPP from 90.05 to 91.72. These results support learning from downstream outcomes as an effective way to exploit draft–target complementarity within speculative decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.