acceptodds
Under review as a conference paper at ICLR 2027

Architecture-Conditioned KV Compressibility

Abstract

KV-cache compression is usually evaluated as a deployment method, but a prior question is often left implicit: which native cache representations contain structure worth compressing? We study KV compressibility as a conditional property of representation, architecture, context length, domain, and tensor folding. The main audited screen contains nine completed instruction-tuned models from Qwen2.5, Llama, and Mistral, spanning 0.5–22.2B parameters, four domains, four contexts from 512 to 32K, and three seeds: 432 valid cells and 1,296 representation rows. A separate exact-length matrix contains 420 cells from three models, four domains, seven contexts from 512 to 32K, and five seeds. Three results emerge. First, representation ordering is universal in both protocols: pre-RoPE keys beat post-RoPE keys, which beat values, in 432/432 and 420/420 paired cells. Second, quantized tensor-train (QTT) compatibility is architecture-conditioned and non-monotone: Qwen2.5-7B reaches a 7.32× median pre-RoPE-key QTT ratio in the main screen, while neighboring Qwen2.5-3B and 14B are near 1×; the protocol-separated Qwen2.5-1.5B exact-context matrix reaches 26.82×. Third, long context amplifies representation mismatch: from 512 to 32K, rank-16 error rises by 0.0246 for pre-RoPE K, 0.1448 for post-RoPE K, and 0.0819 for V. The paper provides a structural screening prior, not a claim of resident-memory or latency improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.