SpecGrove: Reusable Tree Capacity for Concurrent Diffusion Speculation
Abstract
Concurrent block-diffusion speculative decoding must share limited verification capacity across requests with heterogeneous returns from tree expansion. We introduce SpecGrove, which turns a single DFlash proposal and DDTree expansion per request into reusable, nested capacity tiers. SpecGrove scores tiers by covered draft-prefix mass and jointly selects heterogeneous tree sizes using measured packed-pass cost. An exact-row Bellman frontier optimizes the resulting finite-tier surrogate, and selected trees share one request-isolated target pass. We establish optimality for this finite-tier objective and characterize conditions under which packed verification preserves the joint autoregressive output law. Across Qwen3-4B and Qwen3-8B on four reasoning and code-generation benchmarks, SpecGrove improves complete wall-clock throughput by 13.0–16.8% over matched DFlash and 4.7–39.0% over capacity-feasible fixed-size DDTree at concurrency eight. Under Poisson arrivals, it reduces tail latency relative to matched allocation controls.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.