acceptodds
Under review as a conference paper at ICLR 2027

Beyond Jailbreak Rates: Harm Centric Evaluation of Multi Turn Jailbreak Attacks and Defenses

Abstract

Multi-turn jailbreak attacks are the standard probe of safety aligned language models, yet a session is still judged by a binary verdict, the jailbreak rate, which records only whether a breach occurred. A binary verdict cannot tell how much harm was released, how much of the objective remains within the attacker’s reach after a harmless reply, or whether the defense also refuses ordinary requests. We introduce a harm-centric evaluation framework in which every round is graded on realized severity of harm and residual risk, over refusal on a published benign suite is reported beside harm, defenses are ordered on severity accumulated over the session, and residual risk is a forward-looking diagnostic. One primary rater scored 12,000 rounds from fixed ten round Crescendo derived trajectories across six conditions and four open weight model families; two independent raters each rescored a held out 90 round subset and agreed closely on five of the six com- ponents. Evaluation methods are compared on two measures. Resolution is the number of defense pairs an instrument tells apart by their performance. Of 60 pair- wise defense comparisons, the refusal keyword rule separates 10, two published classifiers, HarmBench and Llama Guard 3, are able to distinguish 26 and 20 re- spectively, and the harm-centric score separates 51. Similarity to expert judgment is measured against an independent blinded security expert on a random sample of 200 rounds, and the harm-centric score reaches 95 percent. The classifiers, in turn, call safe the rounds holding 29 and 19 percent of the harm the raters found. We also find that several defense orderings change across target families, while the utility gate verdict depends on the benign suite. Code and scored data are available in an anonymous repository.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.