acceptodds
Under review as a conference paper at ICLR 2027

RelaySpec: Adapting Frozen Speculative Drafters with Linear Maps

Abstract

Speculative decoding accelerates autoregressive LLM inference by using a lightweight drafter to propose tokens that the target verifies in parallel. Recent high-performance drafters improve proposal quality by conditioning on intermediate features from the target model. This coupling, however, makes the drafter difficult to reuse when the target changes. LoRA adaptation shifts these features, and replacing the target can even change their dimensionality. We introduce RelaySpec, a method that adapts a pretrained drafter to a new target by learning one linear map per target feature stream while keeping the drafter's transformer frozen. On Qwen3-4B with GSM8K, KiCad and NanoCoder LoRAs, RelaySpec improves DFlash throughput by 53% on average over the unadapted drafter. When transferring a Qwen3-4B drafter to Qwen3-8B with in-domain fitting data, the adapted drafter reaches 94-100% of native-drafter throughput across math, code, and chat. The reverse transfer, from Qwen3-8B to Qwen3-4B, reaches 88-97%. A single fit on mixed data reaches 99.4% of native throughput on Spec-Bench. In vLLM serving, three LoRA-fine-tuned targets share a single frozen drafter, achieving 77% higher aggregate throughput than the unadapted drafter at 32 concurrent clients. RelaySpec also generalizes to EAGLE-3, improving throughput by up to 105% over the unadapted drafter across three LoRA-fine-tuned targets. It also outperforms SD by 3.6-10.2% even when SD also fine-tunes the drafter, while training 19 fewer parameters, and outperforms OmniDraft by 1.9-52.4%. RelaySpec adapts an existing drafter with thousands of target-generated trajectories instead of training a new one from scratch, which for DFlash takes roughly 800,000 examples.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.