acceptodds
Under review as a conference paper at ICLR 2027

ReTree: Budget-Efficient Tree Speculative Decoding via Path Guidance and Sibling Recovery

Abstract

Tree speculative decoding spends a fixed target-model budget on multiple candidate continuations, but exact verification can stop while scored siblings remain unused. We present ReTree, an approximate decoding policy that recovers an existing sibling using historical divergence counts, current target probability support, and stop-token constraints. Recovery reuses the same tree forward pass and introduces no learned correction head. A request-local path prior additionally changes which candidates enter the tree. On Qwen3-4B and Qwen3-8B at a 16-node budget, the measured speed gains over DDTree are 4.2–5.1% at temperature 0 and 8.8–9.1% at temperature 1. Recovery accounts for most of the gain; path guidance increases committed length but does not consistently improve speed over recovery alone. Additional CSD-style and MARS-style comparisons show that other relaxed verifiers can exceed ReTree's speed at some operating points. These comparisons do not establish quality-matched superiority. We characterize the contribution as tree-local reuse of already scored candidates, with an explicit speed–quality trade-off rather than preservation of the target distribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.