Accelerating Suffix-Based Jailbreak Evaluation with Prefix-Shared KV Cache
Abstract
Suffix jailbreak attacks are a systematic tool for red-teaming large language models (LLMs), but they require evaluating many candidate suffixes before finding an effective jailbreak. This paper presents Prefix-Shared KV Cache (PSKV), a plug-and-play inference optimization tailored to suffix jailbreak evaluation. PSKV exploits the structure that many candidate prompts share the same harmful instruction prefix while differing only in the candidate suffix. Instead of redundantly recomputing or physically duplicating the prefix cache, PSKV stores one compact prefix KV cache and expands it lazily at each layer during candidate evaluation. This design preserves PyTorch autograd compatibility and enables larger batched searches with lower memory overhead. Across six suffix attacks and five open-source LLMs, PSKV reduces inference time by about 40% and peak memory by about 50% while preserving attack effectiveness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.