acceptodds
Under review as a conference paper at ICLR 2027

RELAY: Momentum Drafting from Hidden Transitions for Speculative Decoding

Abstract

Speculative decoding accelerates large language model inference by using a lightweight drafter to propose tokens for parallel verification. Recent feature-guided drafters improve draft quality by reusing hidden features from the target forward pass, but these features are typically treated as static endpoints, exposing what the target model has computed without explicitly modeling how its representations evolve toward prediction. This leaves the lightweight drafter to infer the missing transformation, making farther-ahead predictions harder to align with the target model. We instead observe that transitions between target-model stages provide a direct drafting signal, as they capture representation changes rather than merely their resulting states. Based on this, we propose RELAY, a lightweight drafter that performs momentum drafting from hidden transitions. RELAY decomposes the target forward pass into low-, mid-, and top-stage features, converts their stagewise differences into early and late momentum states, and recurrently refines these momenta with an Attention-MLP architecture during draft-time generation. By propagating hidden-transition momentum across draft steps, RELAY uses target-side computation as evolving guidance. Experiments across multiple backbones, benchmarks, and serving settings show consistent efficiency gains, achieving up to 31.3% higher speedup and 30.0% longer average acceptance length over EAGLE-3. Code is available at https://anonymous.4open.science/r/RELAY-8421/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.