iKVC: Inference-Time KV-Cache Correction for Quantized LLMs in Long-Horizon Tasks
Abstract
Post-training quantization (PTQ) reduces the memory footprint of large language models (LLMs) and can improve inference throughput. However, the resulting quantization errors can significantly degrade accuracy, creating a fundamental trade-off between the efficiency gained from quantization and the accuracy retained from the full-precision model, particularly in long-horizon generation. Prior work mitigates this degradation by preserving the KV cache of selected prompt tokens at full precision, thereby reducing errors introduced during prefill. This strategy, however, does not address errors introduced during decoding, as quantized weights continuously perturb the key and value representations of newly generated tokens, and these errors accumulate in the KV cache, altering subsequent predictions. We empirically observe that this error accumulation becomes non-negligible as a large number of tokens are generated, and we theoretically derive that the accumulated KV-cache errors can cause the attention output to deviate from the values computed by the original full-precision model. Motivated by this finding, we propose inference-time KV-cache Correction (iKVC), a training-free inference method that periodically corrects the KV cache for both prompt and generated tokens before decoding errors substantially accumulate. The correction interval is configurable, allowing iKVC to explicitly control the accuracy–throughput trade-off. Across experiments, iKVC consistently improves the trade-off over PTQ alone in long-horizon tasks. Our results identify KV-cache error accumulation as an important, previously underexplored factor in achieving accurate and efficient long-horizon tasks with quantized LLMs, suggesting that managing such errors during inference can complement PTQ techniques.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.