acceptodds
Under review as a conference paper at ICLR 2027

A Sampling-Centric Framework for Tree Speculative Decoding

Abstract

Tree speculative decoding accelerates large language model inference by drafting and verifying multiple candidate continuations in parallel. Top-candidate dynamic-tree methods such as EAGLE-2/3 follow a topology-first paradigm: they expand deterministic candidates and globally prune them by score, so draft–target overlap outside the retained set cannot contribute to acceptance. Introducing sampled candidates into this pipeline creates a selection problem: under global pruning, a sample's realized identity can affect its own survival. Existing sampling extensions such as RheoSampling retain expand-then-prune and therefore must explicitly control how samples are scored and pruned to preserve losslessness. We introduce Solvo, a sampling-centric framework that replaces expand-then-prune with allocate-then-sample. At each layer, a global path-score threshold commits slots to parents before any token is drawn, and slots are never reranked afterward, so a sampled token can condition subsequent expansion but not its own retention. Under stochastic decoding, each non-root candidate group places top-ranked deterministic tokens in every slot except the last, which is filled with a sample from the remaining draft mass. Using a martingale argument over tree growth, we prove that Solvo preserves the exact target-model distribution under adaptive growth and cost-aware stopping. Across five target models and six tasks, Solvo improves end-to-end throughput by 7.3-10.2% over EAGLE-3 under stochastic decoding, and still by 2.9-7.8% under greedy decoding, where sampling plays no role.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.