Don’t Compress Away the Evidence: Secure Provenance Auditing for LLM KV Caches
Abstract
Efficient KV cache compression is essential for deploying long-context large language models, but it introduces an unexamined conflict with model ownership auditing. Existing fingerprinting methods embed ownership evidence through trigger patterns that are deliberately positioned away from routine task semantics—precisely the kind of patterns that cache compression discards as “unimportant". Therefore, we propose Fingerprint-Preserving Causal KV Cache Compression (FinCa), a framework that treats ownership evidence and task utility as joint constraints. Our key insight is to replace static trigger patterns with a low-rank perturbation in a KV projection that generates an input-conditioned structured signal in the runtime cache. Complementing this generation mechanism, we reserve fingerprint-bearing coordinates during compression and apply quantization-index modulation to protect them, ensuring that the ownership signal survives cache reconstruction. As a result, FinCa embeds a robust fingerprint signal from only 20 documents and achieves signal recovery in most evaluated settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.