acceptodds
Under review as a conference paper at ICLR 2027

CHARD: Efficient Block-Diffusion Speculative Decoding via Chunked Verification and Immediate Redrafting

Abstract

Block-diffusion drafters generate multiple consecutive draft tokens in a single forward pass and often yield long accepted prefixes. In batched serving, however, requests within a physical target batch typically share a common verification width. This coupling creates two inefficiencies. First, when a request rejects before the end of the verification window, target computation on the remaining suffix cannot contribute to its output. Second, accepted progress varies substantially across concurrent requests and workloads. Although requests undergo the same fixed-width verification round, low-acceptance requests commit fewer tokens, reducing their token production rate and extending the tail of time per output token (TPOT). Consequently, long accepted prefixes do not fully translate into serving gains. We present CHARD (Chunk-Aware Redrafting and Decoding), a lossless serving system that decouples proposal horizon, verification granularity, and request recovery. Chunked verification consumes a long proposal in bounded-width chunks and reuses the remaining suffix only when the current chunk is fully accepted and the target-produced boundary token matches the stored draft anchor. This avoids target computation on invalidated suffixes. Immediate redrafting resolves request state after every chunk: a terminated request becomes eligible for a replacement proposal at the next scheduling opportunity, while requests with valid suffixes continue verification. We implement CHARD in SGLang and evaluate it on Qwen3-8B and Qwen3.5-27B with their corresponding DFlash drafters. For 16-token proposals, CHARD improves throughput by up to 30.4% and reduces request-level p99 TPOT by up to 27.4% relative to one-shot DFlash with the same proposal horizon. CHARD retains long draft proposals, requires no learned controller, and preserves the target model's deterministic decoding semantics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.