acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Target: From Imitation to Collaboration in Speculative Decoding

Abstract

Speculative decoding accelerates large language model inference by letting a small draft model propose tokens for parallel verification by a larger target. For tasks evaluated by answer correctness, however, agreement with the target is an imperfect criterion: a weaker draft can occasionally offer a continuation that leads to a better final answer. Exploiting these disagreements requires learning from delayed outcomes while retaining the efficiency of parallel verification. We present Collaborative Speculative Decoding (CoSpec), a framework that learns arbitration within speculative decoding from the outcomes of complete rollouts. An arbitrator jointly observes the context, draft block, and target verification tokens through a hybrid attention mask, and decides whether to retain the draft or fall back to the target at each mismatch. Reinforcement learning combines final-answer correctness with accepted-length shaping to train these decisions. CoSpec retains parallel target verification while optimizing task utility rather than preserving the target distribution. We evaluate two model families on reasoning, code, and knowledge benchmarks. With LLaMA-3.3-70B as the target, CoSpec achieves a mean speedup of 3.96× over target-only decoding and raises the mean score on GSM8K, HumanEval, and MBPP from 90.05 to 91.72. These results support learning from downstream outcomes as an effective way to exploit draft–target complementarity within speculative decoding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.