acceptodds
Under review as a conference paper at ICLR 2027

STEP: Separating Trajectory Evolution and Prediction for Speculative Decoding

Abstract

Speculative decoding accelerates large language model inference, but its efficiency depends on sustaining accurate draft predictions over multiple future positions. Autoregressive drafters such as EAGLE-3 use a shared hidden state for token prediction and recurrent state propagation. We observe that high token prediction accuracy can coexist with substantial recurrent-state deviation from target-model features, motivating separate representations for these two roles. We introduce STEP, a feature–logit disentangled drafter that separates trajectory evolution from token prediction. To stabilize recurrent propagation, we employ a staged feature-alignment curriculum that uses target-model fused features as an early training reference before gradually removing the auxiliary constraint. We further exploit the resulting long-horizon drafting capability with confidence-guided progressive tree expansion, selectively extending promising branches to obtain longer accepted continuations. Across four target models, five benchmarks, and two temperatures, STEP outperforms EAGLE-3 in all 40 settings, with average relative gains of 10.1% in decoding speed and 15.1% in acceptance length. Ablations support complementary roles for disentanglement and alignment, while progressive tree expansion further improves decoding efficiency. In SGLang, STEP also achieves 2.9–11.9% higher throughput than EAGLE-3 across all tested batch sizes under the same speculative token budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.