acceptodds
Under review as a conference paper at ICLR 2027

Co-Evolving Trees and Control Policies: Bridging Posterior Optimality and Online Efficiency in Speculative Decoding

Abstract

Tree-based speculative decoding accelerates LLM inference through parallel verification of candidate continuations. Posterior verification records and a calibrated latency model enable statistically optimized tree portfolios with near-optimal surrogate throughput in a finite candidate space. Pre-drafting routing, however, commits an entire cycle with limited information, making mistaken template choices costly and hindering realization of posterior gains. Our preliminary experiments identify two advantages of decisions during drafting: richer observations and reuse of shared computation across alternative actions. We therefore formulate an ideal joint objective for draft structures and online policies: expected committed tokens per expected complete-cycle time. We propose Oracle-Guided Structure–Policy Optimization (OGSPO), using a common net-gain objective for coupled conditional updates. It first solves for posterior-optimal decision depths using verification records and modeled costs in a mixed-integer fractional program, then alternates policy and structure updates. With the draft structure fixed, posterior records and a latency model estimate every legal terminal action’s net gain for policy learning. Sampled same-trace executions correct policy-improvement estimates. With the policy response fixed, state-conditioned reach and stop probabilities weight mixed-integer optimization of the shared nested draft structure. We establish conditional unbiasedness of policy-improvement estimates under the sampling protocol and finite-domain optimality bounds for the structural proxy under a fixed response. Across twelve configurations spanning three models, four devices, and two temperatures, OGSPO improves four-workload mean complete-request throughput over EAGLE-3 by 9.2% on average and up to 19.3%. Comparisons and ablations further support the benefits of joint structure–policy design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.