HARP: Horizon-Aware Representation Preference for Parallel Speculative Decoding
Abstract
Speculative decoding (SD) accelerates large language model (LLM) inference by letting a lightweight drafter propose tokens that the target model verifies in parallel. Parallel drafters push this paradigm further by predicting an entire block of future tokens from a single target computation at the verified prefix, yet they treat all positions in the block uniformly, fusing target hidden states into one shared representation and ranking candidates by token confidence. We identify prediction horizon, the distance of a position from the verified prefix, as a key factor in both drafting and verification. On the drafting side, the target layer with the strongest predictive signal shifts systematically from upper layers toward intermediate layers as the horizon grows, so different future positions favor different target representations. On the verification side, we prove that for repair trees preserving the main draft chain, the additional accepted length decomposes exactly at the first mismatch into prefix survival, alternative-token probability, and recoverable continuation, showing that the value of a repair branch depends on more than per-token confidence. Based on these results, we propose HARP (Horizon-Aware Representation Preference), a training-free inference framework that assigns each horizon its preferred combination of already-computed target layers and allocates repair branches by this reachable value. On Qwen3-4B with a frozen DFlash drafter, HARP improves end-to-end throughput by 19.7% over chain verification, compared with 9.7% for confidence-based allocation, and achieves higher throughput than recent draft-tree methods while verifying fewer nodes. A factorial study shows that the two components provide complementary gains, and HARP improves throughput by 10.8–14.3% across six target–drafter pairs. Prediction horizon thus provides a principled axis for exploiting target representations and verification structure in training-free SD acceleration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.