acceptodds
Under review as a conference paper at ICLR 2027

BitLadder: Bit-Level Selective Loading of the KV Cache for Efficient Long-Context LLM Inference

Abstract

Long-context large language models (LLMs) repeatedly load a growing key-value (KV) cache during autoregressive decoding, making attention memory-bandwidth bound. Existing methods reduce this traffic through query-aware sparse attention or KV-cache quantization, but neither aligns loading with current token importance: sparse attention uses a coarse keep-or-skip mask that cannot differentiate retained contributions and may discard the aggregate contribution of long-tail tokens, whereas quantization uses static bit-width assignments that cannot track query-dependent importance. Our analysis shows that token importance is graded and changes across decoding steps. We present BitLadder, an algorithm-system co-designed framework for bit-level selective loading that dynamically adapts loaded KV-cache bit widths to token importance at each decoding step. To support dynamic bit widths without storing multiple quantized copies or performing per-step requantization, we propose truncation-consistent nested quantization (TCNQuant), which encodes the KV-cache once and derives lower-bit representations by truncating the same code. To translate fewer loaded bits into practical speedup, we further design PlaneWeave, a memory layout that stores this nested representation so that a GPU kernel loads only the bits each token needs and supports efficient mixed-precision execution. We evaluate BitLadder on long-context understanding (LongBenchV2, MRCR) and reasoning (GPQA, MATH500, AIME25) benchmarks. At matched accuracy, BitLadder achieves 3.0× and 2.3× speedups over full attention and the state-of-the-art baseline, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.