KVRecover: Cross-Model Prompt Recovery from Leaked KV Caches
Abstract
Key–value (KV) caches retain input information that can be exploited to reconstruct private prompts.However, reconstruction attacks that rely on a large source model can require substantial GPU resources.We introduce KVRecover, a cross-model prompt-reconstruction method that enables a smaller frozen receiver to recover input content from complete leaked KV caches.Its model-pair-specific converter uses depth-dependent projection to map source states into the receiver's representation space, while local cross-layer fusion coordinates the converted states.Progressive supervision combines cache alignment, attention alignment, and text reconstruction to optimize the converted caches for input recovery.Once trained, the converter can be distributed and reused on unseen prompts within the same model pair, without source-model inference or further converter optimization during recovery.Experiments across three Qwen3 model pairs and four datasets achieve a macro-average BERTScore F1 of up to 0.834.Recovery from Qwen3-235B-A22B caches using Qwen3-14B further demonstrates the feasibility of the attack across a substantial model-size gap.These findings show that leaked caches can remain vulnerable even to attackers unable to deploy the source model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.