DStreamer: Acceptance-Aligned Diffusion Streaming for High-Throughput Serving and Online RLVR
Abstract
Long-form generation is increasingly the dominant cost in high-throughput LLM serving and online reinforcement learning with verifiable rewards (RLVR). Speculative decoding increases token progress per target-model invocation by verifying multiple proposals in parallel, but at throughput-oriented batch sizes its effectiveness depends on keeping proposal and state overhead low as concurrency and context grow. Online RLVR adds a distinct challenge: as the target policy changes during training, a drafter calibrated once can become progressively stale. We introduce DStreamer, a lossless speculative decoding framework with a single-pass diffusion-based in-model adapter that reuses the target model's cached KV state and requires no separate draft KV cache. DStreamer warm-starts the adapter with consistency distillation, then optimizes proposal paths with an acceptance-aligned reinforcement-learning objective under autoregressive verification, directly targeting the prefix-sensitive criterion that determines speculative utility. We evaluate DStreamer in the throughput-oriented regimes where speculative decoding is most constrained by proposal and state overhead: continuous-batched serving, long-form generation, and synchronous online RLVR. Across Qwen and GPT-OSS models, DStreamer expands the serving frontier under concurrencies 64-1024, outperforming DFlash and EAGLE-3 throughout the evaluated range and reaching up to higher aggregate throughput than autoregressive decoding at matched per-request rate, while requiring no separate draft KV cache. DStreamer++ extends acceptance-oriented training to evolving policies while preserving exact on-policy rollouts, achieving - end-to-end RLVR speedup across Qwen3-4B, 8B, and 14B under synchronous H100 training, while preserving final policy quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.