Beyond Jailbreak Rates: Harm Centric Evaluation of Multi Turn Jailbreak Attacks and Defenses
Abstract
Multi-turn jailbreak attacks are the standard probe of safety aligned language models, yet a session is still judged by a binary verdict, the jailbreak rate, which records only whether a breach occurred. A binary verdict cannot tell how much harm was released, how much of the objective remains within the attacker’s reach after a harmless reply, or whether the defense also refuses ordinary requests. We introduce a harm-centric evaluation framework in which every round is graded on realized severity of harm and residual risk, over refusal on a published benign suite is reported beside harm, defenses are ordered on severity accumulated over the session, and residual risk is a forward-looking diagnostic. One primary rater scored 12,000 rounds from fixed ten round Crescendo derived trajectories across six conditions and four open weight model families; two independent raters each rescored a held out 90 round subset and agreed closely on five of the six com- ponents. Evaluation methods are compared on two measures. Resolution is the number of defense pairs an instrument tells apart by their performance. Of 60 pair- wise defense comparisons, the refusal keyword rule separates 10, two published classifiers, HarmBench and Llama Guard 3, are able to distinguish 26 and 20 re- spectively, and the harm-centric score separates 51. Similarity to expert judgment is measured against an independent blinded security expert on a random sample of 200 rounds, and the harm-centric score reaches 95 percent. The classifiers, in turn, call safe the rounds holding 29 and 19 percent of the harm the raters found. We also find that several defense orderings change across target families, while the utility gate verdict depends on the benign suite. Code and scored data are available in an anonymous repository.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.