acceptodds
Under review as a conference paper at ICLR 2027

Model Priming: Unverified Handoffs for Efficient Decoding

Abstract

The fast adoption of large language models is straining global compute, making inference efficiency both an economic and environmental imperative. We investigate primed decoding: a scheduled handoff policy in which a small model M_q continues a generation initiated by a larger, more capable model M_p. Unlike speculative decoding, primed decoding performs no verification and requires no router, retraining, or shared serving infrastructure. We evaluate it on five model pairs from two model families and six reasoning, mathematics, and code benchmarks, measuring accuracy together with throughput, cost, and energy on H200 GPUs. For instance, generating only the first 1024 tokens with Qwen3.5-122B-A10B and the rest with Qwen3.5-4B recovers 92% of Qwen3.5-122B-A10B's performance and reduces single-request latency by 1.6–1.8× on long generations. At the same time, it cuts cost by 60% and energy consumption by 68%. These results suggest that a short high-capability prefix can establish a reasoning trajectory that a smaller model often sustains, enabling substantial offloading of generation to cheaper models without token-level coordination.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.