STEP: Separating Trajectory Evolution and Prediction for Speculative Decoding
Abstract
Speculative decoding accelerates large language model inference, but its efficiency depends on sustaining accurate draft predictions over multiple future positions. Autoregressive drafters such as EAGLE-3 use a shared hidden state for token prediction and recurrent state propagation. We observe that high token prediction accuracy can coexist with substantial recurrent-state deviation from target-model features, motivating separate representations for these two roles. We introduce STEP, a feature–logit disentangled drafter that separates trajectory evolution from token prediction. To stabilize recurrent propagation, we employ a staged feature-alignment curriculum that uses target-model fused features as an early training reference before gradually removing the auxiliary constraint. We further exploit the resulting long-horizon drafting capability with confidence-guided progressive tree expansion, selectively extending promising branches to obtain longer accepted continuations. Across four target models, five benchmarks, and two temperatures, STEP outperforms EAGLE-3 in all 40 settings, with average relative gains of 10.1% in decoding speed and 15.1% in acceptance length. Ablations support complementary roles for disentanglement and alignment, while progressive tree expansion further improves decoding efficiency. In SGLang, STEP also achieves 2.9–11.9% higher throughput than EAGLE-3 across all tested batch sizes under the same speculative token budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.