DSpark-OPD: Acceptance-Aligned On-Policy Distillation for Semi-Autoregressive Speculative Drafting
Abstract
Standard rejection-sampling speculative decoding preserves the target distribution for any proposal model, but its speed depends on the length of the draft prefix that survives verification. Semi-autoregressive drafters face two training mismatches: offline training observes target-generated prefixes rather than draft-induced states, and tokenwise objectives do not reflect that an early rejection invalidates the remaining suffix. We present DSpark-OPD, an on-policy post-training method for Markov DSpark. During rollout, the frozen target verifies samples from the deployed Markov-corrected proposal. During replay, target and drafter distributions are evaluated on the same sampled predecessor chain, and optimization is applied to the corrected proposal used by the verifier. We introduce Survival Loss, a differentiable pathwise surrogate based on target–draft overlap that assigns earlier transitions credit for later positions that remain reachable. Across Qwen3-8B and Qwen3-14B, adding Survival Loss to the matched DSpark-OPD-Base objective improves every evaluated dataset–temperature comparison. Under greedy decoding, public-average engine acceptance length rises by 4.7% at 8B and 5.4% at 14B relative to DSpark-OPD-Base; relative to the released DSpark checkpoints, the complete pipeline gains are 12.8% and 6.0%. At concurrency 8, public-benchmark throughput improves by 4.9% and 6.1% over DSpark-OPD-Base, reaching 2.07× and 2.40× target-only throughput. Controlled Qwen3-8B ablations show that Survival Loss remains competitive with local and fixed-decay variants at block length 7 while providing state-adaptive prefix credit without a hand-designed decay schedule.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.