acceptodds
Under review as a conference paper at ICLR 2027

SCD: Enabling Speculative Overlap for Target-Conditioned Draft Models

Abstract

A stronger drafter can improve acceptance in speculative decoding, but its extra computation may offset the gain. Target–draft disaggregation places the models on separate devices, yet each request still waits for verification before the next draft. Speculative speculative decoding (SSD) predicts the next starting point to overlap these steps, but target-conditioned drafters must also wait for target hidden states. We propose Staged Conditional Drafting (SCD) to shorten this wait. The first draft layers use earlier target features to compute predicted next-round branches during verification. The remaining layers reuse these intermediate activations and continue with deeper target features. For batched requests, a small drafter handles prediction misses in parallel. We implement SCD with Two-stage DSpark and KDA + Target-KV. Both preserve accepted length on a 4B target; the KDA version also does so on 14B. With the same models, batch, and two-GPU layout, using target features earlier reduces total decode time by 15.47%. Under greedy decoding in native SGLang, two-GPU overlap improves batch-size-one decode rate by 6.86–9.32% over single-GPU serial execution across two draft architectures and three tasks. It also improves throughput in fully occupied batched rounds by 8.50–17.99% over single-GPU DSpark. KDA + Target-KV accommodates 13% more concurrent sessions within the same GPU memory budget, at a 2.17% lower single-GPU decode rate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.