acceptodds
Under review as a conference paper at ICLR 2027

ByteBudget: Optimizing Weight–KV precision allocation for CPU inference

Abstract

Autoregressive decoding on CPUs is memory-bandwidth-bound; generating each token requires streaming model weights, whose traffic grows with the context length. However, the quantization methods typically optimize weights or KV caches separately, and a fixed precision policy cannot account for changes in context length, hardware characteristics, or kernel efficiency. We introduce ByteBudget, a hardware-and-context-aware planner that converts a decode-latency target into a calibrated effective-traffic budget and jointly allocates precision across weight blocks and a KV cache to minimize estimated quality loss. The framework consists of three components: a format-aware CPU cost model, modular quality-loss curves supplied by off-the-shelf quantizers, and a discrete budgeted optimizer. For practical deployment, we use an efficient marginal-gain planner to produce mixed-precision weight and KV cache configurations, with a dynamic-programming oracle and lagrangian lower bound to audit solution quality. Since KV cache traffic grows with context length while weight traffic remains roughly constant, ByteBudget adapts precision allocation changing deployment conditions. Our framework integrates the quantization methods with CPU-based LLM decoding kernels and allocates precision across layers to satisfy deployment-specific latency constraints, rather than a uniform bit-width.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.