Elastic Decoding: Controlling LLM Inference Speed via Contextual Sparsity
Abstract
Large language model (LLM) inference in memory-bound regimes is constrained by a fundamental trade-off between generation quality and decoding speed. In this regime, contextual sparsity (e.g., mixtures of experts) accelerates autoregressive generation by reducing the number of active parameters loaded per token, while preserving the model's total parameter count. We introduce Elastic Decoding, a contextual sparsity method that can be applied post-hoc to any trained LLM, dense or MoE, returns a single checkpoint with controllable sparsity levels, requires less than 5B finetuning tokens, and is at the Pareto frontier of speed and accuracy. To achieve this, we frame contextual sparsification as a multiple-choice knapsack problem with bytes-per-token as the budget, allocating capacity heterogeneously across the architecture before distilling the model into a single sparse checkpoint. Elastic Decoding can reduce the active number of parameters by up to 83% and, paired with optimized on-device kernels, achieves end-to-end speed-ups of up to 4.6x (e.g., on Gemma4-31B). Our sparsified Qwen3-14B runs faster and more accurately than Qwen3-8B and Qwen3-4B. Importantly, the quality-speed trade-off can be controlled during decoding in real-time, on a token-by-token basis, with negligible switching overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.