Speculative Fine-Tuning: Accelerated Supervised Fine-Tuning with Speculative Heads
Abstract
Supervised fine-tuning is a standard method for adapting large language models to downstream tasks, but it can still require substantial compute, especially when repeated across models, datasets, and deployment settings. In this work, we show that speculative decoding, originally introduced for inference acceleration, can also be used to reduce the number of optimization steps required for supervised fine-tuning. Our key observation is that gradients from speculative heads are strongly correlated with gradients from the main language-model head. We use this correlation to construct control-variate gradient estimators that reduce stochastic gradient variance while preserving the corresponding training objective. We further show that the standard multi-token prediction (MTP) objective admits a natural control-variate interpretation, connecting MTP training to variance reduction. Moreover, we find that the correlation varies across the layers of the model, allowing us to develop a layer-aware variant of the method. Across supervised fine-tuning experiments, our proposed optimizer requires up to **85%** fewer training samples, corresponding to a **6.7×** improvement in sample efficiency, and implies an estimated **2.65×** wall-clock speedup using measured per-step latency, while reducing the final main loss by up to **12.2%**. To our knowledge, this is the first work to use speculative decoding for optimizer updates, separating the optimization role of speculative heads from their conventional use as auxiliary prediction heads. Across five evaluation seeds, MTP+CV improves GSM8K accuracy by 3.08 and 2.98 percentage points on Qwen3.5-4B and Gemma-4 E2B, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.