acceptodds
Under review as a conference paper at ICLR 2027

LoadMatch: Balancing Memory and Compute Loads in Quantized Models using Speculative Decoding

Abstract

Quantization has become a standard practice for efficient large language model (LLM) inference. However, the exact mechanism behind this acceleration is often misunderstood. A common misunderstanding is that quantized weights directly speed up arithmetic operations: in reality, low-bit GEMM still incurs higher-precision dequantization, rescaling, and accumulation, so its compute-side speedup is much smaller than its memory-side gain. Consequently, the computation time itself remains largely unchanged. However, the perceived speedup stems from the efficiency gained by reducing the transfer time of quantized weights from the GPU memory to its cache. Because of this, quantized LLMs achieve significant speedups only in memory-bound scenarios, whereas their improvements in compute-bound regimes are marginal. To truly maximize the decoding throughput of quantized LLMs, a system must achieve a load-matched state where the memory transfer duration and computation time are equal so that neither the memory bus nor the compute units remain idle. To reach this state, we propose LoadMatch, which balances memory and computation costs using speculative decoding. We leverage speculative decoding, in which a draft model predicts multiple future tokens, called candidate tokens, and the target model verifies them in parallel. By precisely controlling the candidate-token count processed at once, our method deliberately matches the memory and computation workload. Empirically, LoadMatch compounds the speedups from quantization and speculative decoding, delivering a measured end-to-end speedup of over FP16 autoregressive greedy decoding on standard benchmarks. We will release the project upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.