Request Forest: Jointly Growing Verification Trees for Speculative Decoding
Abstract
Block-parallel draft models can predict multiple future tokens in a single forward pass, yet the benefits of speculative decoding under high concurrency depend on how limited verification resources are allocated. Collapsing candidates into a single path and assigning a fixed verification window to each request limits the ability to accommodate differences in draft quality across requests and exploit alternative branches within each request, while increasing intermediate verification-state overhead in hybrid architectures. We propose Request Forest (RF), which jointly grows verification trees across requests under a shared budget. RF reuses the candidates and predecessor-conditioned scores provided by a block-parallel drafter, allowing eligible nodes across requests to compete for the budget according to their path priorities. This process jointly determines tree size, depth, and branching structure, yielding a prefix-closed verification forest. RF then packs the forest into a compact, branch-isolated layout, verifies it in a single target-model forward pass, and commits only the states along accepted paths. Theoretical analysis establishes that, under the specified scoring assumptions and budget constraints, this growth strategy maximizes the sum of priorities of selected candidates. Experiments on four hybrid-architecture models and four benchmarks demonstrate the effectiveness of RF. On Qwen3.8-27B at concurrency levels of 64–160, RF improves geometric-mean throughput by 56.6% over DFlash and 13.9% over DFlash2. At a concurrency level of 128, it reduces GPU Memory reserved for intermediate recurrent verification states per tensor-parallel rank by 50.4% relative to DFlash2.These results show that shared-budget tree growth effectively translates diffusion draft diversity into throughput gains under high concurrency.Code is available at https://anonymous.4open.science/r/Request-Forest-2844/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.