acceptodds
Under review as a conference paper at ICLR 2027

BagDraft: Next-Bag-of-Tokens Drafting for Lossless Speculative Decoding

Abstract

The speedup of speculative decoding is fundamentally bounded by draft token throughput, yet existing drafters either incur sequential passes (EAGLE), dilute model capacity across disjoint heads (Medusa), or tie capacity to rigid per-slot block predictions (DFlash). We propose BagDraft, which reformulates speculative drafting as a next-bag-of-tokens prediction task. Crucially, the informational debt of discarding sequence order requires no parametric recovery in speculative decoding, as exact target verification guarantees lossless generation. Operating via a single forward pass with a shared logit vector, BagDraft directly models future -token sets under a multi-hot cross-entropy objective, using discrete anchor blocks to ensure strict input isomorphism between training and inference. At inference, we mitigate the permutation explosion by formalizing path probabilities via Plackett–Luce sampling without replacement; its proven prefix-monotonicity property mathematically justifies expectation-based pruning to construct compact permutation trees, verified by the target model in a single forward pass. Across 7 benchmarks on Qwen3-4B/8B, BagDraft achieves an average acceptance length of and a wall-clock speedup on 4B ( over DFlash, over EAGLE-3), proving that speculative drafters can surpass state-of-the-art efficiency without dedicated per-position heads or autoregressive rollout, relying solely on an offset-decayed shared distribution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.