Model Priming: Unverified Handoffs for Efficient Decoding
Abstract
The fast adoption of large language models is straining global compute, making inference efficiency both an economic and environmental imperative. We investigate primed decoding: a scheduled handoff policy in which a small model M_q continues a generation initiated by a larger, more capable model M_p. Unlike speculative decoding, primed decoding performs no verification and requires no router, retraining, or shared serving infrastructure. We evaluate it on five model pairs from two model families and six reasoning, mathematics, and code benchmarks, measuring accuracy together with throughput, cost, and energy on H200 GPUs. For instance, generating only the first 1024 tokens with Qwen3.5-122B-A10B and the rest with Qwen3.5-4B recovers 92% of Qwen3.5-122B-A10B's performance and reduces single-request latency by 1.6–1.8× on long generations. At the same time, it cuts cost by 60% and energy consumption by 68%. These results suggest that a short high-capability prefix can establish a reasoning trajectory that a smaller model often sustains, enabling substantial offloading of generation to cheaper models without token-level coordination.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.