acceptodds
Under review as a conference paper at ICLR 2027

KV Compression Cannot Double as a Jailbreak Defense: The Harmfulness-then-Refusal Trajectory Guides Token Deletion

Abstract

KV-cache compression has recently been repurposed as a jailbreak defense, on the assumption that attention-importance rankings and key-space geometry can contribute to a safety decision. We show that safety and KV compression do not couple. We find that (i) attention-based token importance is uncorrelated with harmfulness and with refusal, so eviction ordered by it cannot localize the harmful request, and (ii) the key projection carries no direction that mediates refusal, since keys route attention but do not determine the content that is refused. We identify a Harmfulness-then-Refusal Trajectory, in which the harmfulness signal peaks at the request and the refusal signal immediately after it; the model identifies the request as harmful before it constructs a refusal. We therefore propose HaRD (Harmfulness-Refusal-guided token Deletion), a training-free defense spanning three major families of jailbreak attack, semantic manipulation, automated adversarial search and context accumulation, which deletes tokens before prefill and leaves the model and its inference unchanged. HaRD reads harmfulness and refusal directions to detect the jailbreak and locate the harmful query, then deletes the jailbreak prompt around it, exposing the query and leaving the model's own alignment to refuse. Experiments across seven jailbreak attacks show that our method cuts the aggregate harm rate from 69% to 8% and raises explicit refusal from 30% to 88%, at minimal utility loss and no over-refusal cost, and with substantially reduced latency and memory on attack prompts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.