acceptodds
Under review as a conference paper at ICLR 2027

LMTC: Quantizing KV Caches for How Attention Reads Them in Omni-Modal LLMs

Abstract

Omni-modal Large Language Models (OmniLLMs) unify the understanding of video, audio, and text, yet their continuous streaming inputs generate an overwhelming volume of tokens, making the Key-Value (KV) cache a major memory bottleneck. Existing KV cache quantization methods are suboptimal for OmniLLMs because they either minimize simple reconstruction error or rely on query activations that are unavailable when streaming tokens are cached, ignoring both the attention reading mechanism and cross-modal heterogeneity. To address these challenges, we propose Logit-Metric Transform Coding (LMTC), a training-free and query-free KV cache quantization framework that minimizes the expected attention error using only the query projection weights. Specifically, LMTC quantizes Keys using a Modality-Conditioned Transform Coder guided by a Query-Free Logit Metric derived directly from query projection weights, thereby encoding each modality in its own basis and allocating precision via reverse water-filling to the directions attention reads. Simultaneously, LMTC handles Values through Phase-Offset Rounding, a metadata-free mechanism that offsets the rounding grids of spatio-temporally adjacent visual tokens so that rounding errors cancel out during attention aggregation rather than accumulate. Extensive evaluations across three OmniLLMs on eight benchmarks show that LMTC is virtually lossless at 3 effective bits per scalar, retaining over 99% of BF16 average accuracy. Compared to BF16 FlashDecoding-v2, LMTC achieves up to a decoding speedup, a KV cache memory reduction, and a throughput boost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.