acceptodds
Under review as a conference paper at ICLR 2027

EchoSpec: Efficient Target-Integrated Speculative Decoding via Output-Projection Adaptation and Post-Target Refinement

Abstract

Autoregressive decoding is a major bottleneck in large language model inference due to its sequential and memory-bandwidth-bound nature. Speculative decoding alleviates this bottleneck by drafting multiple future tokens for parallel verification. While recent methods improve draft quality using separate auxiliary transformers, these models introduce target-specific parameters and significant training costs. Furthermore, the auxiliary weights and KV caches increase serving memory, particularly at high concurrency. In this paper, we introduce EchoSpec, a target-integrated speculative decoding framework that reuses the frozen target model for drafting with minimal adaptation. At each layer, we apply a learned bias exclusively to the attention output projection at draft positions. This approach is highly parameter-efficient compared to broadly adapting the target model. Nonetheless, drafting in parallel with verification prevents the draft positions from accessing the bonus token. To resolve the resulting information gap, we introduce a lightweight post-target refinement block in EchoSpec rather than scaling up target-side adaptation. As a result, EchoSpec requires only 2.2-4.5K target-generated training samples, features 27.9-32.1x fewer trainable parameters than DFlash, and eliminates the need for a separate draft KV cache. Under long-context, high-concurrency workloads, it reduces the active memory footprint and delivers higher output throughput than separate transformer drafters such as DFlash and EAGLE-3.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.