Why Freeze the Target Model? Aligning the LLM Toward Its Small Drafter for Speculative Decoding
Abstract
Speculative decoding accelerates large language model (LLM) inference by using a small language model (SLM) to draft several tokens that the LLM verifies in parallel. Existing methods typically focus on training the SLM towards LLM to fully imitate the LLM’s distribution. However, we observe that, in open-ended generation, many rejected SLM outputs are semantically correct but differ from the LLM in wording or style. For a capacity-limited SLM, matching the LLM’s wording and style in addition to generating correct content remains challenging. This raises a key question: can the LLM instead accommodate the SLM’s expression style while retaining its own capabilities? Building on this idea, we propose SpecL2S, a two-stage mutual alignment framework. We first use content-aware distillation to train the SLM to prioritize content accuracy over stylistic imitation. We then freeze the trained SLM and optimize a lightweight soft prompt for the LLM, adapting its expression preferences to the SLM while preserving its content-critical behavior. Together, these two stages align the SLM toward the LLM’s capabilities and the LLM toward the SLM’s expression style. The resulting pair uses standard parallel verification without an additional semantic judge. Across five open-ended benchmarks, SpecL2S increases the mean number of accepted tokens per target call from 3.05 to 4.57 and achieves a 2.43× decoding speedup, while quality remains nearly unchanged
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.